Skip to main navigation Skip to search Skip to main content

Cross-Modal Mixup Enhance Foundation Model Adaptation for Few-Shot Learning

  • Jiuqian Dai
  • , Zhenyan Ji*
  • , Zechang Xiong
  • , Jiqiang Liu
  • , Shen Yin
  • , Huihui Wang
  • *Corresponding author for this work
  • Beijing Jiaotong University
  • Norwegian University of Science and Technology
  • Northeastern University China

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

In few-shot learning, adapting foundation models has emerged as a powerful approach to bridge the gap between limited task-specific data and the rich knowledge encoded in pre-trained models. Existing adaptation methods based on CLIP primarily focus on improving the network architecture to enhance the learning ability, which has been proven to cause overfitting and reduce generalizability, but overlook the influence of data augmentation. In this work, we first analyze the differences between our work and previous work from the perspective of probability theory. Then, we concentrate on data augmentation and propose innovative Cross-Modal Mixup methods designed to enhance the model's cross-modal reasoning and classification capabilities, as well as its overall generalization. Specifically, based on dimensions match or not, the Cross-Modal Mixup methods comprise two forms: Cross Mixup (CM) for data augmentation with features of the same dimension, and Cross Expand Mixup (CEM) with features of different dimensions. Extensive experiments have demonstrated the effectiveness of our CM and CEM.

Original languageEnglish
Title of host publicationProceedings - 2025 IEEE Smart World Congress, SWC 2025, 2025 IEEE Ubiquitous Intelligence and Computing, Autonomous and Trusted Computing, Digital Twin, Metaverse, Scalable Computing and Communications
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages1094-1099
Number of pages6
ISBN (Electronic)9798331575984
DOIs
StatePublished - 2025
Externally publishedYes
Event2025 IEEE Smart World Congress, SWC 2025 - Calgary, Canada
Duration: 18 Aug 202522 Aug 2025

Publication series

NameProceedings - 2025 IEEE Smart World Congress, SWC 2025, 2025 IEEE Ubiquitous Intelligence and Computing, Autonomous and Trusted Computing, Digital Twin, Metaverse, Scalable Computing and Communications

Conference

Conference2025 IEEE Smart World Congress, SWC 2025
Country/TerritoryCanada
CityCalgary
Period18/08/2522/08/25

Keywords

  • CLIP
  • Data augmentation
  • Few-shot learning
  • Multimodal
  • Parameter efficient finetune

Fingerprint

Dive into the research topics of 'Cross-Modal Mixup Enhance Foundation Model Adaptation for Few-Shot Learning'. Together they form a unique fingerprint.

Cite this