Skip to main navigation Skip to search Skip to main content

Multimodal intent recognition based on text-guided cross-modal attention

  • Zhengyi Li
  • , Junjie Peng*
  • , Xuanchao Lin
  • , Zesu Cai
  • *Corresponding author for this work
  • Shanghai University
  • School of Computer Science and Technology, Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

In natural language understanding, intent recognition stands out as a crucial task that has drawn significant attention. While previous research focuses on intent recognition using task-specific unimodal data, real-world scenarios often involve human intents expressed through various ways, including speech, tone of voice, facial expressions, and actions. This prompts research into integrating multimodal information to more accurately identify human intent. However, existing intent recognition studies often fuse textual and non-textual modalities without considering their quality gap. The gap in feature quality across different modalities hinders the improvement of the model’s performance. To address this challenge, we propose a multimodal intent recognition model to enhance non-textual modality features. Specifically, we enrich the semantics of non-textual modalities by replacing redundant information through text-guided cross-modal attention. Additionally, we introduce a text-centric adaptive fusion gating mechanism to capitalize on the primary role of text modality in intent recognition. Extensive experiments on two multimodal task datasets show that our proposed model performs better in all metrics than state-of-the-art multimodal models. The results demonstrate that our model efficiently enhances non-textual modality features and fuses multimodal information, showing promising potential for intent recognition.

Original languageEnglish
Article number690
JournalApplied Intelligence
Volume55
Issue number7
DOIs
StatePublished - May 2025
Externally publishedYes

Keywords

  • Attention mechanism
  • Feature enhancement
  • Intent recognition
  • Multimodal fusion

Fingerprint

Dive into the research topics of 'Multimodal intent recognition based on text-guided cross-modal attention'. Together they form a unique fingerprint.

Cite this