Skip to main navigation Skip to search Skip to main content

Text-guided multimodal depression detection via cross-modal feature reconstruction and decomposition

  • Ziqiang Chen
  • , Dandan Wang
  • , Liangliang Lou
  • , Shiqing Zhang*
  • , Xiaoming Zhao
  • , Shuqiang Jiang
  • , Jun Yu
  • , Jun Xiao
  • *Corresponding author for this work
  • TaiZhou University
  • Zhejiang Sci-Tech University
  • CAS - Institute of Computing Technology
  • Harbin Institute of Technology Shenzhen
  • Zhejiang University

Research output: Contribution to journalArticlepeer-review

Abstract

Depression, a widespread and debilitating mental health disorder, requires early detection to facilitate effective intervention. Automated depression detection integrating audio with text modalities is a challenging yet significant issue due to the information redundancy and inter-modal heterogeneity across modalities. Prior works usually fail to fully learn the interaction of audio–text modalities for depression detection in an explicit manner. To address these issues, this work proposes a novel text-guided multimdoal depression detection method based on a cross-modal feature reconstruction and decomposition framework. The proposed method takes the text modality as the core modality to guide the model to reconstruct comprehensive audio features for cross-modal feature decomposition tasks. Moreover, the designed cross-modal feature reconstruction and decomposition framework aims to disentangle the shared and private features from the text-guided reconstructed comprehensive audio features for subsequent multimodal fusion. Besides, a bi-directional cross-attention module is designed to interactively learn simultaneous and mutual correlations across modalities for feature enhancement. Extensive experiments are performed on the DAIC-WoZ and E-DAIC datasets, and the results show the superiority of the proposed method on multimodal depression detection tasks, outperforming the state-of-the-arts.

Original languageEnglish
Article number102861
JournalInformation Fusion
Volume117
DOIs
StatePublished - May 2025
Externally publishedYes

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 3 - Good Health and Well-being
    SDG 3 Good Health and Well-being

Keywords

  • Cross-modal feature reconstruction
  • Depression detection
  • Feature decomposition
  • Multimodal fusion

Fingerprint

Dive into the research topics of 'Text-guided multimodal depression detection via cross-modal feature reconstruction and decomposition'. Together they form a unique fingerprint.

Cite this