Skip to main navigation Skip to search Skip to main content

ASA: An Auditory Spatial Attention Dataset with Multiple Speaking Locations

  • Zijie Lin
  • , Tianyu He
  • , Siqi Cai*
  • , Haizhou Li
  • *Corresponding author for this work
  • The Chinese University of Hong Kong, Shenzhen
  • Shenzhen Research Institute of Big Data
  • National University of Singapore
  • University of Bremen

Research output: Contribution to journalConference articlepeer-review

Abstract

Recent studies have demonstrated the feasibility of localizing an attended sound source from electroencephalography (EEG) signals in a cocktail party scenario. This is referred to as EEG-enabled Auditory Spatial Attention Detection (ASAD). Despite the promise, there is a lack of ASAD datasets. Most existing ASAD datasets are recorded from two speaking locations. To bridge this gap, we introduce a new Auditory Spatial Attention (ASA) dataset, featuring multiple speaking locations of sound sources. The new dataset is designed to challenge and refine deep neural network solutions in real-world applications. Furthermore, we build a channel attention convolutional neural network (CA-CNN) as a reference model for ASA, that serves as a competitive benchmark for future studies.

Original languageEnglish
Pages (from-to)437-441
Number of pages5
JournalProceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
DOIs
StatePublished - 2024
Externally publishedYes
Event25th Interspeech Conferece 2024 - Kos Island, Greece
Duration: 1 Sep 20245 Sep 2024

Keywords

  • Auditory spatial attention
  • EEG
  • channel attention
  • cocktail party problem

Fingerprint

Dive into the research topics of 'ASA: An Auditory Spatial Attention Dataset with Multiple Speaking Locations'. Together they form a unique fingerprint.

Cite this