Skip to main navigation Skip to search Skip to main content

Visual question answering for bridge damage inspection using a multi-modal large language model

  • School of Transportation Science and Engineering, Harbin Institute of Technology
  • National University of Singapore

Research output: Contribution to journalArticlepeer-review

Abstract

Current artificial intelligence approaches for bridge inspection mainly focus on isolated tasks, such as damage classification, localization, or image captioning, thereby lacking the flexibility to provide the comprehensive, multi-modal assessments required for practical bridge maintenance. To address this limitation, this paper proposes a Bridge Inspection Visual Question Answering (BIVQA) model built upon a multi-modal large language model. Unlike conventional models that provide static outputs, BIVQA interprets natural language queries to generate context-aware responses and simultaneously localizes damage. The architecture comprises three components: a vision encoder for high-level feature extraction, a multi-modal fusion encoder that integrates visual and textual embeddings, and a decoder that generates natural language answers while leveraging multi-head cross-attention for precise damage localization. Experimental results validate the robustness of BIVQA, demonstrating its capability to accurately perform component identification, damage description, and localization all within a unified framework, offering a more interactive and interpretable solution for automated bridge inspection.

Original languageEnglish
Article number107125
JournalAutomation in Construction
Volume190
DOIs
StatePublished - Oct 2026
Externally publishedYes

Keywords

  • Bridge inspection
  • Multi-modal large language model
  • Visual question answering

Fingerprint

Dive into the research topics of 'Visual question answering for bridge damage inspection using a multi-modal large language model'. Together they form a unique fingerprint.

Cite this