Abstract
The integration of computer vision and multimodal images has been a focus for evaluating road condition, but adaptability and cross-modal fusion issues cause pretrained models to perform poorly in complex scenarios. In this study, a novel intelligent recognition framework integrating image fusion, object identification, and visualization representation was proposed for pavement distress detection. The framework integrates multiple optimization modules including Depthwise Separable Convolution (DWConv), Simple Attention Mechanism (SimAM), Wise Intersection over Union (Wise-IoU) and Gradient-weighted Class Activation Mapping (Grad-CAM), which enables it to prioritize contextual semantic information and facilitate precise feature extraction in multimodal pavement imagery, thereby enhancing both model interpretability and detection accuracy for small-scale objects. To verify its effectiveness, the proposed algorithm is comprehensively evaluated on the different datasets, and results show macro-averaged classification scores of 80.2% and 58.7% on self-constructed and public datasets, respectively, outperforming other models. Ablation studies further demonstrate the excellence of module improvement in enhancing model performance. This research provides valuable reference for multimodal pavement detection systems in complex environments.
| Original language | English |
|---|---|
| Article number | 115292 |
| Journal | Engineering Applications of Artificial Intelligence |
| Volume | 179 |
| DOIs | |
| State | Published - 1 Sep 2026 |
| Externally published | Yes |
Keywords
- Image fusion
- Infrared thermography
- Pavement distress detection
- Visualization representation
- You only look once algorithm
Fingerprint
Dive into the research topics of 'Pavement distress detection using multimodal image fusion and enhanced convolutional neural network in complex scenarios'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver