Abstract
In order to learn more emotionally inclined video and speech representations through auxiliary tasks, and improve the effect of multi-modal fusion, this paper proposes a multi-modal sentiment recognition method based on multi-task learning. A multimodal sharing layer is used to learn the sentiment information of the visual and acoustic modes. The experiment on MOSI and MOSEI data sets shows that adding two auxiliary single-modal sentiment recognition tasks can learn more effective single-modal sentiment representations, and improve the accuracy of sentiment recognition by 0.8% and 2.5% respectively.
| Translated title of the contribution | A Multi-modal Sentiment Recognition Method Based on Multi-task Learning |
|---|---|
| Original language | Chinese (Traditional) |
| Pages (from-to) | 7-15 |
| Number of pages | 9 |
| Journal | Beijing Daxue Xuebao (Ziran Kexue Ban)/Acta Scientiarum Naturalium Universitatis Pekinensis |
| Volume | 57 |
| Issue number | 1 |
| DOIs | |
| State | Published - 20 Jan 2021 |
| Externally published | Yes |
Fingerprint
Dive into the research topics of 'A Multi-modal Sentiment Recognition Method Based on Multi-task Learning'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver