Skip to main navigation Skip to search Skip to main content

An End-to-End Transformer with Progressive Tri-Modal Attention for Multi-modal Emotion Recognition

  • Harbin Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Recent works on multi-modal emotion recognition move towards end-to-end models, which can extract the task-specific features supervised by the target task compared with the two-phase pipeline. In this paper, we propose a novel multi-modal end-to-end transformer for emotion recognition, which can effectively model the tri-modal features interaction among the textual, acoustic, and visual modalities at the low-level and high-level. At the low-level, we propose the progressive tri-modal attention, which can model the tri-modal feature interactions by adopting a two-pass strategy and can further leverage such interactions to significantly reduce the computation and memory complexity through reducing the input token length. At the high-level, we introduce the tri-modal feature fusion layer to explicitly aggregate the semantic representations of three modalities. The experimental results on the CMU-MOSEI and IEMOCAP datasets show that ME2ET achieves the state-of-the-art performance. The further in-depth analysis demonstrates the effectiveness, efficiency, and interpretability of the proposed tri-modal attention, which can help our model to achieve better performance while significantly reducing the computation and memory cost (Our code is available at https://github.com/SCIR-MSA-Team/UFMAC.).

Original languageEnglish
Title of host publicationPattern Recognition and Computer Vision - 6th Chinese Conference, PRCV 2023, Proceedings
EditorsQingshan Liu, Hanzi Wang, Rongrong Ji, Zhanyu Ma, Weishi Zheng, Hongbin Zha, Xilin Chen, Liang Wang
PublisherSpringer Science and Business Media Deutschland GmbH
Pages396-408
Number of pages13
ISBN (Print)9789819985395
DOIs
StatePublished - 2024
Event6th Chinese Conference on Pattern Recognition and Computer Vision, PRCV 2023 - Xiamen, China
Duration: 13 Oct 202315 Oct 2023

Publication series

NameLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume14431 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference6th Chinese Conference on Pattern Recognition and Computer Vision, PRCV 2023
Country/TerritoryChina
CityXiamen
Period13/10/2315/10/23

Keywords

  • Feature fusion
  • Multi-modal emotion recognition
  • Multi-modal transformer

Fingerprint

Dive into the research topics of 'An End-to-End Transformer with Progressive Tri-Modal Attention for Multi-modal Emotion Recognition'. Together they form a unique fingerprint.

Cite this