Skip to main navigation Skip to search Skip to main content

Bridge damage description using adaptive attention-based image captioning

  • School of Transportation Science and Engineering, Harbin Institute of Technology
  • School of Civil Engineering, Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Current many vision-based research uses various classification, detection, and segmentation methods to identify bridge damage. Instead of these numerical results, a highly abstract natural language description is a more suitable method to summarize and transmit bridge inspection processes and results to humans. This paper presents an end-to-end image captioning-based bridge damage comprehensive description network (BDCD-Net) for describing and locating bridge damage. BDCD-Net consists of two parts: an image feature encoder (extracting multi-level image features from bridge damage images) and a bridge damage description generation decoder (employing an adaptive attention mechanism to selectively utilize image features to generate descriptions and locate damage). The descriptions include component types, damage categories, relative spatial positions of the damaged components and bridges, and shooting angles of the image. The effectiveness of BDCD-Net was validated using images collected from real bridges with annotated descriptions. The results indicate the significant potential of fully automated bridge inspection.

Original languageEnglish
Article number105525
JournalAutomation in Construction
Volume165
DOIs
StatePublished - Sep 2024
Externally publishedYes

Keywords

  • Adaptive attention
  • Bridge damage
  • Comprehensive description
  • Image captioning
  • Multimodal learning

Fingerprint

Dive into the research topics of 'Bridge damage description using adaptive attention-based image captioning'. Together they form a unique fingerprint.

Cite this