TY - GEN
T1 - Independent Feature Enhanced Crossmodal Fusion for Match-Mismatch Classification of Speech Stimulus and EEG Response
AU - Fan, Shitong
AU - Wang, Wenbo
AU - Xiao, Feiyang
AU - Zhang, Shiheng
AU - Zhu, Qiaoxi
AU - Guan, Jian
N1 - Publisher Copyright:
©2024 IEEE.
PY - 2024
Y1 - 2024
N2 - It is crucial for auditory attention decoding to classify matched and mismatched speech stimuli with corresponding EEG responses by exploring their relationship. However, existing methods often adopt two independent networks to encode speech stimulus and EEG response, which neglect the relationship between these signals from the two modalities. In this paper, we propose an independent feature enhanced crossmodal fusion model (IFE-CF) for match-mismatch classification, which leverages the fusion feature of the speech stimulus and the EEG response to achieve auditory EEG decoding. Specifically, our IFE-CF contains a crossmodal encoder to encode the speech stimulus and the EEG response with a two-branch structure connected via crossmodal attention mechanism in the encoding process, a multi-channel fusion module to fuse features of two modalities by aggregating the interaction feature obtained from the crossmodal encoder and the independent feature obtained from the speech stimulus and EEG response, and a predictor to give the matching result. In addition, the causal mask is introduced to consider the time delay of the speech-EEG pair in the crossmodal encoder, which further enhances the feature representation for match-mismatch classification. Experiments demonstrate our method’s effectiveness with better classification accuracy, as compared with the baseline of the Auditory EEG Decoding Challenge 2023.
AB - It is crucial for auditory attention decoding to classify matched and mismatched speech stimuli with corresponding EEG responses by exploring their relationship. However, existing methods often adopt two independent networks to encode speech stimulus and EEG response, which neglect the relationship between these signals from the two modalities. In this paper, we propose an independent feature enhanced crossmodal fusion model (IFE-CF) for match-mismatch classification, which leverages the fusion feature of the speech stimulus and the EEG response to achieve auditory EEG decoding. Specifically, our IFE-CF contains a crossmodal encoder to encode the speech stimulus and the EEG response with a two-branch structure connected via crossmodal attention mechanism in the encoding process, a multi-channel fusion module to fuse features of two modalities by aggregating the interaction feature obtained from the crossmodal encoder and the independent feature obtained from the speech stimulus and EEG response, and a predictor to give the matching result. In addition, the causal mask is introduced to consider the time delay of the speech-EEG pair in the crossmodal encoder, which further enhances the feature representation for match-mismatch classification. Experiments demonstrate our method’s effectiveness with better classification accuracy, as compared with the baseline of the Auditory EEG Decoding Challenge 2023.
KW - Auditory EEG decoding
KW - cross-attention
KW - feature fusion
KW - multi-modal learning
UR - https://www.scopus.com/pages/publications/85216395006
U2 - 10.1109/ISCSLP63861.2024.10800649
DO - 10.1109/ISCSLP63861.2024.10800649
M3 - 会议稿件
AN - SCOPUS:85216395006
T3 - 2024 14th International Symposium on Chinese Spoken Language Processing, ISCSLP 2024
SP - 209
EP - 213
BT - 2024 14th International Symposium on Chinese Spoken Language Processing, ISCSLP 2024
A2 - Qian, Yanmin
A2 - Jin, Qin
A2 - Ou, Zhijian
A2 - Ling, Zhenhua
A2 - Wu, Zhiyong
A2 - Li, Ya
A2 - Xie, Lei
A2 - Tao, Jianhua
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 14th International Symposium on Chinese Spoken Language Processing, ISCSLP 2024
Y2 - 7 November 2024 through 10 November 2024
ER -