Skip to main navigation Skip to search Skip to main content

Enhancing Audio Retrieval with Attention-based Encoder for Audio Feature Representation

  • Feiyang Xiao
  • , Qiaoxi Zhu
  • , Jian Guan*
  • , Wenwu Wang
  • *Corresponding author for this work
  • College of Computer Science and Technology, Harbin Engineering University
  • University of Technology Sydney
  • University of Surrey

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Pretrained audio neural networks (PANNs) has been successful in a range of machine audition applications. But its limitation in recognising relationships between acoustic scenes and events impacts its performance in language-based audio retrieval, which retrieves audio signals from a dataset based on natural language textual queries. This paper proposes the attention-based audio encoder to exploit contextual associations between acoustic scenes/events, using self-attention or graph attention with different loss functions for language-based audio retrieval. Our experimental results show that the proposed attention-based method outperforms most of state-of-the-art methods, with self-attention performing better than graph attention. In addition, the selection of different loss functions (i.e., NT-Xent loss or supervised contrastive loss) does not have as significant an impact on the results as the selection of the attention strategy.

Original languageEnglish
Title of host publication31st European Signal Processing Conference, EUSIPCO 2023 - Proceedings
PublisherEuropean Signal Processing Conference, EUSIPCO
Pages755-759
Number of pages5
ISBN (Electronic)9789464593600
DOIs
StatePublished - 2023
Externally publishedYes
Event31st European Signal Processing Conference, EUSIPCO 2023 - Helsinki, Finland
Duration: 4 Sep 20238 Sep 2023

Publication series

NameEuropean Signal Processing Conference
ISSN (Electronic)2076-1465

Conference

Conference31st European Signal Processing Conference, EUSIPCO 2023
Country/TerritoryFinland
CityHelsinki
Period4/09/238/09/23

Keywords

  • Language-based audio retrieval
  • attention mechanism
  • audio representation
  • multimodal learning

Fingerprint

Dive into the research topics of 'Enhancing Audio Retrieval with Attention-based Encoder for Audio Feature Representation'. Together they form a unique fingerprint.

Cite this