Abstract
Current structural damage segmentation models are often trained based on substantial pixel-level labels for specific structural components and damage types. To address this issue, this paper establishes a transformer-based large vision model for universal structural damage segmentation, incorporating a pre-trained transformer-based frozen backbone and a fine-tuned CNN-based segmentation head. A synthetic loss function of correlation loss and contrastive loss is proposed. A self-supervised correlation learning procedure is designed to ensure cross-level feature alignment. The contrastive loss across student-teacher networks is designed to learn intra-instance similarity and inter-instance separability. A contrastive learning strategy is employed to fine-tune the segmentation head by exponential moving average with momentum updating. The proposed method is validated on a multi-scale image dataset for cable-supported bridges, concrete bridges, and post-earthquake buildings. The recognition accuracy, generalization ability, robustness under complex background, and superiority to conventional supervised and unsupervised segmentation models are demonstrated.
| Original language | English |
|---|---|
| Article number | 106256 |
| Journal | Automation in Construction |
| Volume | 176 |
| DOIs | |
| State | Published - Aug 2025 |
Keywords
- Cross-level feature correlation alignment
- Large vision model
- Self-supervised learning
- Teacher-student contrastive learning
- Universal structural damage segmentation
Fingerprint
Dive into the research topics of 'Transformer-based large vision model for universal structural damage segmentation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver