Abstract
Auditory Attention Detection (AAD) seeks to identify the target attended by a listener in a multi-speaker environment using electroencephalography (EEG) signals. A critical yet under-explored aspect of this task is the choice of speech representations that model the relationship between the auditory stimulus and the elicited neural signals. In this study, we conduct a comprehensive and comparative analysis of four distinct speech representations for AAD, namely speech envelope, spectral features (STFT), a deep acoustic representation (Wav2Vec), and a semantic representation that reflects linguistic content. To our knowledge, this is the first study to integrate a semantic speech representation for AAD. We further introduce a cross-modal feature fusion architecture designed to model the complex dependencies between these speech features and the corresponding EEG dynamics. Evaluation on three publicly available datasets demonstrates that the semantic representation consistently and substantially outperforms all other representations, achieving performance gains of 3.74%, 1.12%, and 2.17% over the envelope, STFT, and deep acoustic representation baselines, respectively. This finding provides a new perspective for cross-modal neural decoding, highlighting the importance of semantic information in both human and machine auditory attention. Code will be available at: https://github.com/Lindahahaha/AAD-speech-representation/
| Original language | English |
|---|---|
| Pages (from-to) | 146-151 |
| Number of pages | 6 |
| Journal | Pattern Recognition Letters |
| Volume | 203 |
| DOIs | |
| State | Published - May 2026 |
| Externally published | Yes |
Keywords
- Auditory attention
- Cocktail party problem
- EEG
- Speech representation
Fingerprint
Dive into the research topics of 'The effect of speech representations on EEG-based auditory attention detection'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver