TY - GEN
T1 - Multimodal Dialog System
T2 - 29th ACM International Conference on Multimedia, MM 2021
AU - Zhang, Haoyu
AU - Liu, Meng
AU - Gao, Zan
AU - Lei, Xiaoqiang
AU - Wang, Yinglong
AU - Nie, Liqiang
N1 - Publisher Copyright:
© 2021 ACM.
PY - 2021/10/17
Y1 - 2021/10/17
N2 - Multimodal dialog system has attracted increasing attention from both academia and industry over recent years. Although existing methods have achieved some progress, they are still confronted with challenges in the aspect of question understanding (i.e., user intention comprehension). In this paper, we present a relational graph-based context-aware question understanding scheme, which enhances the user intention comprehension from local to global. Specifically, we first utilize multiple attribute matrices as the guidance information to fully exploit the product-related keywords from each textual sentence, strengthening the local representation of user intentions. Afterwards, we design a sparse graph attention network to adaptively aggregate effective context information for each utterance, completely understanding the user intentions from a global perspective. Moreover, extensive experiments over a benchmark dataset show the superiority of our model compared with several state-of-the-art baselines.
AB - Multimodal dialog system has attracted increasing attention from both academia and industry over recent years. Although existing methods have achieved some progress, they are still confronted with challenges in the aspect of question understanding (i.e., user intention comprehension). In this paper, we present a relational graph-based context-aware question understanding scheme, which enhances the user intention comprehension from local to global. Specifically, we first utilize multiple attribute matrices as the guidance information to fully exploit the product-related keywords from each textual sentence, strengthening the local representation of user intentions. Afterwards, we design a sparse graph attention network to adaptively aggregate effective context information for each utterance, completely understanding the user intentions from a global perspective. Moreover, extensive experiments over a benchmark dataset show the superiority of our model compared with several state-of-the-art baselines.
KW - attribute-enhanced text representation
KW - multimodal dialog system
KW - sparse relational context modeling
UR - https://www.scopus.com/pages/publications/85117621163
U2 - 10.1145/3474085.3475234
DO - 10.1145/3474085.3475234
M3 - 会议稿件
AN - SCOPUS:85117621163
T3 - MM 2021 - Proceedings of the 29th ACM International Conference on Multimedia
SP - 695
EP - 703
BT - MM 2021 - Proceedings of the 29th ACM International Conference on Multimedia
PB - Association for Computing Machinery, Inc
Y2 - 20 October 2021 through 24 October 2021
ER -