Skip to main navigation Skip to search Skip to main content

Efficient matching of Transformer-enhanced features for accurate vision-based displacement measurement

  • Haoyu Zhang
  • , Stephen Wu
  • , Xiangyun Luo
  • , Yong Huang*
  • , Hui Li
  • *Corresponding author for this work
  • Harbin Institute of Technology
  • School of Civil Engineering, Harbin Institute of Technology
  • Institute of Statistical Mathematics
  • The Graduate University for Advanced Studies

Research output: Contribution to journalArticlepeer-review

Abstract

Computer vision technology and monitoring videos have been employed to obtain structural displacement measurements. Noniterative algorithms are mainly designed for rapid tracking of the motions of individual image points, rather than dense motion fields. Iterative algorithms are limited to estimating motion fields with small amplitudes and require high computation cost to achieve high accuracy. This paper introduces a noniterative method for vision-based measurements that balances speed and density. The method employs an attention-based matching strategy applied to Transformer-enhanced image features. Motion priors and a physics-informed denoising approach are integrated to improve measurement accuracy. Tested on challenging truss and cable-stayed bridge vibration videos, the method demonstrated superior displacement measurement performance compared to conventional approaches. It also achieved greater robustness to brightness changes and partial occlusions while requiring minimal human intervention. This method supports the development of automated and affordable vibration monitoring systems.

Original languageEnglish
Article number105962
JournalAutomation in Construction
Volume171
DOIs
StatePublished - Mar 2025

Keywords

  • Attention mechanism
  • Computer vision
  • Displacement measurement
  • Subpixel feature matching
  • Transformer
  • Vibration monitoring

Fingerprint

Dive into the research topics of 'Efficient matching of Transformer-enhanced features for accurate vision-based displacement measurement'. Together they form a unique fingerprint.

Cite this