Abstract
Current artificial intelligence approaches for bridge inspection mainly focus on isolated tasks, such as damage classification, localization, or image captioning, thereby lacking the flexibility to provide the comprehensive, multi-modal assessments required for practical bridge maintenance. To address this limitation, this paper proposes a Bridge Inspection Visual Question Answering (BIVQA) model built upon a multi-modal large language model. Unlike conventional models that provide static outputs, BIVQA interprets natural language queries to generate context-aware responses and simultaneously localizes damage. The architecture comprises three components: a vision encoder for high-level feature extraction, a multi-modal fusion encoder that integrates visual and textual embeddings, and a decoder that generates natural language answers while leveraging multi-head cross-attention for precise damage localization. Experimental results validate the robustness of BIVQA, demonstrating its capability to accurately perform component identification, damage description, and localization all within a unified framework, offering a more interactive and interpretable solution for automated bridge inspection.
| Original language | English |
|---|---|
| Article number | 107125 |
| Journal | Automation in Construction |
| Volume | 190 |
| DOIs | |
| State | Published - Oct 2026 |
| Externally published | Yes |
Keywords
- Bridge inspection
- Multi-modal large language model
- Visual question answering
Fingerprint
Dive into the research topics of 'Visual question answering for bridge damage inspection using a multi-modal large language model'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver