Skip to main navigation Skip to search Skip to main content

Self-adaptive image-text fusion for medical image classification

  • Jian Zhang
  • , Kaihao He
  • , Zunlei Feng
  • , Shuifa Sun*
  • , Xiaoyan Sun
  • , Zhenming Yuan
  • , Jun Yu
  • *Corresponding author for this work
  • Hangzhou Normal University
  • Hangzhou Dianzi University
  • Zhejiang University

Research output: Contribution to journalArticlepeer-review

Abstract

Multimodal classification using both medical images and text reports propels the computer aided disease diagnosis. The performance is susceptible to the quality of image-text fusion. Due to the semantic gap and weak correlation between image and text, current image-text fusion approaches cannot achieve satisfactory results. We propose a self-adaptive image-text fusion approach to multimodal medical image classification. We learn a mapping from image to text to achieve semantic alignment that mitigates the inter-modality semantic gap, and estimate a binary correlation mask with Jensen–Shannon Divergence (JSD) loss to retrieve image and text features that have strong correlations to achieve feature alignment. Then, we propose a parameter-free feature fusion method based on a Simplified-Attention mechanism, which queries image features using text features and concatenates the results to achieve computationally efficient feature fusion. We fuse all the image and text features for medical image classification. Experimental results on three datasets reveal that the proposed approach outperforms a group of state-of-the-art methods, and demonstrates superior medical interpretability.

Original languageEnglish
Article number111715
JournalPattern Recognition
Volume167
DOIs
StatePublished - Nov 2025
Externally publishedYes

Keywords

  • Cross-attention
  • Feature fusion
  • Image classification
  • Multimodal fusion

Fingerprint

Dive into the research topics of 'Self-adaptive image-text fusion for medical image classification'. Together they form a unique fingerprint.

Cite this