Abstract
Current many vision-based research uses various classification, detection, and segmentation methods to identify bridge damage. Instead of these numerical results, a highly abstract natural language description is a more suitable method to summarize and transmit bridge inspection processes and results to humans. This paper presents an end-to-end image captioning-based bridge damage comprehensive description network (BDCD-Net) for describing and locating bridge damage. BDCD-Net consists of two parts: an image feature encoder (extracting multi-level image features from bridge damage images) and a bridge damage description generation decoder (employing an adaptive attention mechanism to selectively utilize image features to generate descriptions and locate damage). The descriptions include component types, damage categories, relative spatial positions of the damaged components and bridges, and shooting angles of the image. The effectiveness of BDCD-Net was validated using images collected from real bridges with annotated descriptions. The results indicate the significant potential of fully automated bridge inspection.
| Original language | English |
|---|---|
| Article number | 105525 |
| Journal | Automation in Construction |
| Volume | 165 |
| DOIs | |
| State | Published - Sep 2024 |
| Externally published | Yes |
Keywords
- Adaptive attention
- Bridge damage
- Comprehensive description
- Image captioning
- Multimodal learning
Fingerprint
Dive into the research topics of 'Bridge damage description using adaptive attention-based image captioning'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver