TY - GEN
T1 - Enhancing Audio Retrieval with Attention-based Encoder for Audio Feature Representation
AU - Xiao, Feiyang
AU - Zhu, Qiaoxi
AU - Guan, Jian
AU - Wang, Wenwu
N1 - Publisher Copyright:
© 2023 European Signal Processing Conference, EUSIPCO. All rights reserved.
PY - 2023
Y1 - 2023
N2 - Pretrained audio neural networks (PANNs) has been successful in a range of machine audition applications. But its limitation in recognising relationships between acoustic scenes and events impacts its performance in language-based audio retrieval, which retrieves audio signals from a dataset based on natural language textual queries. This paper proposes the attention-based audio encoder to exploit contextual associations between acoustic scenes/events, using self-attention or graph attention with different loss functions for language-based audio retrieval. Our experimental results show that the proposed attention-based method outperforms most of state-of-the-art methods, with self-attention performing better than graph attention. In addition, the selection of different loss functions (i.e., NT-Xent loss or supervised contrastive loss) does not have as significant an impact on the results as the selection of the attention strategy.
AB - Pretrained audio neural networks (PANNs) has been successful in a range of machine audition applications. But its limitation in recognising relationships between acoustic scenes and events impacts its performance in language-based audio retrieval, which retrieves audio signals from a dataset based on natural language textual queries. This paper proposes the attention-based audio encoder to exploit contextual associations between acoustic scenes/events, using self-attention or graph attention with different loss functions for language-based audio retrieval. Our experimental results show that the proposed attention-based method outperforms most of state-of-the-art methods, with self-attention performing better than graph attention. In addition, the selection of different loss functions (i.e., NT-Xent loss or supervised contrastive loss) does not have as significant an impact on the results as the selection of the attention strategy.
KW - Language-based audio retrieval
KW - attention mechanism
KW - audio representation
KW - multimodal learning
UR - https://www.scopus.com/pages/publications/85178338143
U2 - 10.23919/EUSIPCO58844.2023.10290096
DO - 10.23919/EUSIPCO58844.2023.10290096
M3 - 会议稿件
AN - SCOPUS:85178338143
T3 - European Signal Processing Conference
SP - 755
EP - 759
BT - 31st European Signal Processing Conference, EUSIPCO 2023 - Proceedings
PB - European Signal Processing Conference, EUSIPCO
T2 - 31st European Signal Processing Conference, EUSIPCO 2023
Y2 - 4 September 2023 through 8 September 2023
ER -