Abstract
Accurate and timely comprehension of surface damage is essential for bridge inspection report automatic generation. Current damage identification multimodal methods are capable of generating rich outputs, but their ability to discern key damage remains suboptimal. To address this challenge, this paper proposes a novel “encoder-reviewer-classifier-decoder” image captioning framework. It integrates a multi-step review network to generate attention-focused thought vectors that capture distinct abstract semantic features of the input image, alongside a damage classifier that enforces category constraints on the decoder. Additionally, it introduces a multi-level supervision mechanism (including generative supervision, discriminant supervision and classification-captioning consistency supervision) to regularize the entire model training process. Compared with three common surface damage image captioning models, the proposed method achieved higher BLEU, ROUGE-L, CIDEr scores and accuracy on a bridge damage image dataset containing 2,701 samples. This study provides a feasible implementation path for multimodal data fusion and transformation in the field of bridge detection.
| Original language | English |
|---|---|
| Article number | 104773 |
| Journal | Advanced Engineering Informatics |
| Volume | 74 |
| DOIs | |
| State | Published - Sep 2026 |
| Externally published | Yes |
Keywords
- Bridge damage
- Category constraint
- Image comprehension
- Multimodal learning
Fingerprint
Dive into the research topics of 'Image-based bridge surface damage comprehension using category-constrained captioning network'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver