Skip to main navigation Skip to search Skip to main content

MAMF-Net: Modality-Adaptive Masked Fusion Network for Speech Emotion Recognition

  • Faculty of Computing, Harbin Institute of Technology
  • Taiyuan University of Technology
  • School of Computer Science and Technology, Harbin Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

This paper introduces a novel multimodal emotion recognition model, the Modality-Adaptive Masked Fusion Network (MAMF-Net), designed to mitigate information loss and improve cross-modal alignment during the fusion of speech and text modalities. MAMF-Net employs an audio-guided text encoder to enhance the semantic representation of text by leveraging the temporal resolution and contextual information inherent in speech, thereby ensuring accurate alignment of modal features. Additionally, the model utilizes a modality transfer-based MAE masking strategy, which effectively captures complementary information between modalities by partially masking transferred information, thus improving fusion effectiveness and system stability. The experimental results show that MAMF-Net outperforms existing methods on datasets such as CMU-MOSI and CMU-MOSEI, highlighting its significant potential for multimodal emotion analysis.

Original languageEnglish
Title of host publication2025 IEEE International Conference on Multimedia and Expo
Subtitle of host publicationJourney to the Center of Machine Imagination, ICME 2025 - Conference Proceedings
PublisherIEEE Computer Society
ISBN (Electronic)9798331594954
DOIs
StatePublished - 2025
Externally publishedYes
Event2025 IEEE International Conference on Multimedia and Expo, ICME 2025 - Nantes, France
Duration: 30 Jun 20254 Jul 2025

Publication series

NameProceedings - IEEE International Conference on Multimedia and Expo
ISSN (Print)1945-7871
ISSN (Electronic)1945-788X

Conference

Conference2025 IEEE International Conference on Multimedia and Expo, ICME 2025
Country/TerritoryFrance
CityNantes
Period30/06/254/07/25

Keywords

  • Speech emotion recognition
  • cross modality
  • masked autoencoder
  • multimodal fusion

Fingerprint

Dive into the research topics of 'MAMF-Net: Modality-Adaptive Masked Fusion Network for Speech Emotion Recognition'. Together they form a unique fingerprint.

Cite this