Abstract
Multimodal classification using both medical images and text reports propels the computer aided disease diagnosis. The performance is susceptible to the quality of image-text fusion. Due to the semantic gap and weak correlation between image and text, current image-text fusion approaches cannot achieve satisfactory results. We propose a self-adaptive image-text fusion approach to multimodal medical image classification. We learn a mapping from image to text to achieve semantic alignment that mitigates the inter-modality semantic gap, and estimate a binary correlation mask with Jensen–Shannon Divergence (JSD) loss to retrieve image and text features that have strong correlations to achieve feature alignment. Then, we propose a parameter-free feature fusion method based on a Simplified-Attention mechanism, which queries image features using text features and concatenates the results to achieve computationally efficient feature fusion. We fuse all the image and text features for medical image classification. Experimental results on three datasets reveal that the proposed approach outperforms a group of state-of-the-art methods, and demonstrates superior medical interpretability.
| Original language | English |
|---|---|
| Article number | 111715 |
| Journal | Pattern Recognition |
| Volume | 167 |
| DOIs | |
| State | Published - Nov 2025 |
| Externally published | Yes |
Keywords
- Cross-attention
- Feature fusion
- Image classification
- Multimodal fusion
Fingerprint
Dive into the research topics of 'Self-adaptive image-text fusion for medical image classification'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver