Skip to main navigation Skip to search Skip to main content

Transformer-based large vision model for universal structural damage segmentation

  • School of Civil Engineering, Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Current structural damage segmentation models are often trained based on substantial pixel-level labels for specific structural components and damage types. To address this issue, this paper establishes a transformer-based large vision model for universal structural damage segmentation, incorporating a pre-trained transformer-based frozen backbone and a fine-tuned CNN-based segmentation head. A synthetic loss function of correlation loss and contrastive loss is proposed. A self-supervised correlation learning procedure is designed to ensure cross-level feature alignment. The contrastive loss across student-teacher networks is designed to learn intra-instance similarity and inter-instance separability. A contrastive learning strategy is employed to fine-tune the segmentation head by exponential moving average with momentum updating. The proposed method is validated on a multi-scale image dataset for cable-supported bridges, concrete bridges, and post-earthquake buildings. The recognition accuracy, generalization ability, robustness under complex background, and superiority to conventional supervised and unsupervised segmentation models are demonstrated.

Original languageEnglish
Article number106256
JournalAutomation in Construction
Volume176
DOIs
StatePublished - Aug 2025

Keywords

  • Cross-level feature correlation alignment
  • Large vision model
  • Self-supervised learning
  • Teacher-student contrastive learning
  • Universal structural damage segmentation

Fingerprint

Dive into the research topics of 'Transformer-based large vision model for universal structural damage segmentation'. Together they form a unique fingerprint.

Cite this