TY - GEN
T1 - Cross-Modal Mixup Enhance Foundation Model Adaptation for Few-Shot Learning
AU - Dai, Jiuqian
AU - Ji, Zhenyan
AU - Xiong, Zechang
AU - Liu, Jiqiang
AU - Yin, Shen
AU - Wang, Huihui
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - In few-shot learning, adapting foundation models has emerged as a powerful approach to bridge the gap between limited task-specific data and the rich knowledge encoded in pre-trained models. Existing adaptation methods based on CLIP primarily focus on improving the network architecture to enhance the learning ability, which has been proven to cause overfitting and reduce generalizability, but overlook the influence of data augmentation. In this work, we first analyze the differences between our work and previous work from the perspective of probability theory. Then, we concentrate on data augmentation and propose innovative Cross-Modal Mixup methods designed to enhance the model's cross-modal reasoning and classification capabilities, as well as its overall generalization. Specifically, based on dimensions match or not, the Cross-Modal Mixup methods comprise two forms: Cross Mixup (CM) for data augmentation with features of the same dimension, and Cross Expand Mixup (CEM) with features of different dimensions. Extensive experiments have demonstrated the effectiveness of our CM and CEM.
AB - In few-shot learning, adapting foundation models has emerged as a powerful approach to bridge the gap between limited task-specific data and the rich knowledge encoded in pre-trained models. Existing adaptation methods based on CLIP primarily focus on improving the network architecture to enhance the learning ability, which has been proven to cause overfitting and reduce generalizability, but overlook the influence of data augmentation. In this work, we first analyze the differences between our work and previous work from the perspective of probability theory. Then, we concentrate on data augmentation and propose innovative Cross-Modal Mixup methods designed to enhance the model's cross-modal reasoning and classification capabilities, as well as its overall generalization. Specifically, based on dimensions match or not, the Cross-Modal Mixup methods comprise two forms: Cross Mixup (CM) for data augmentation with features of the same dimension, and Cross Expand Mixup (CEM) with features of different dimensions. Extensive experiments have demonstrated the effectiveness of our CM and CEM.
KW - CLIP
KW - Data augmentation
KW - Few-shot learning
KW - Multimodal
KW - Parameter efficient finetune
UR - https://www.scopus.com/pages/publications/105035829025
U2 - 10.1109/SWC65939.2025.00175
DO - 10.1109/SWC65939.2025.00175
M3 - 会议稿件
AN - SCOPUS:105035829025
T3 - Proceedings - 2025 IEEE Smart World Congress, SWC 2025, 2025 IEEE Ubiquitous Intelligence and Computing, Autonomous and Trusted Computing, Digital Twin, Metaverse, Scalable Computing and Communications
SP - 1094
EP - 1099
BT - Proceedings - 2025 IEEE Smart World Congress, SWC 2025, 2025 IEEE Ubiquitous Intelligence and Computing, Autonomous and Trusted Computing, Digital Twin, Metaverse, Scalable Computing and Communications
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2025 IEEE Smart World Congress, SWC 2025
Y2 - 18 August 2025 through 22 August 2025
ER -