Skip to main navigation Skip to search Skip to main content

Adapter-Enhanced Hierarchical Cross-Modal Pre-Training for Lightweight Medical Report Generation

  • Ting Yu
  • , Wangwen Lu
  • , Yan Yang
  • , Weidong Han
  • , Qingming Huang
  • , Jun Yu
  • , Ke Zhang*
  • *Corresponding author for this work
  • Hangzhou Normal University
  • Hangzhou Dianzi University
  • Zhejiang Cancer Hospital
  • Zhejiang Normal University
  • University of Chinese Academy of Sciences
  • Harbin Institute of Technology Shenzhen

Research output: Contribution to journalArticlepeer-review

Abstract

Automatic medical report generation is an emerging field that aims to transform medical images into descriptive, clinically relevant narratives, potentially reducing the workload for radiologists significantly. Despite substantial progress, the increasing model parameter size and corresponding marginal performance gains have limited further development and application. To address this challenge, we introduce an Adapter-enhanced Hierarchical cross-modal Pre-training (AHP) strategy for lightweight medical report generation. This approach significantly reduces the pre-trained model's parameter size while maintaining superior report generation performance through our proposed spatial adapters. To further address the issue of inadequate representation of visual space details, we employ a convolutional stem combined with hierarchical injectors and extractors, fully integrating with traditional Vision Transformers to achieve more comprehensive visual representations. Additionally, our cross-modal pre-training model effectively handles the inherent complex visual-textual relationships in medical imaging. Extensive experiments on multiple datasets, including IU X-Ray, MIMIC-CXR, and bladder pathology, demonstrate our model's exceptional generalization and transfer performance in downstream medical report generation tasks, highlighting AHP's potential in significantly reducing model parameters while enhancing report generation accuracy and efficiency.

Original languageEnglish
Pages (from-to)5303-5316
Number of pages14
JournalIEEE Journal of Biomedical and Health Informatics
Volume29
Issue number7
DOIs
StatePublished - 2025
Externally publishedYes

Keywords

  • Cross-modal pre-training
  • lightweight model
  • medical report generation
  • multi-task learning

Fingerprint

Dive into the research topics of 'Adapter-Enhanced Hierarchical Cross-Modal Pre-Training for Lightweight Medical Report Generation'. Together they form a unique fingerprint.

Cite this