Abstract
Models based on deep learning achieve remarkable results in multi-channel speech enhancement. However, they require many parameters and computational resources, challenging practical applications. This paper integrates the idea of branch splitting and channel shuffling with Fourier convolution, proposing a lightweight shuffle Fourier attention network named SF-AN. In addition, we propose flexible channel attention, multi-scale spatial attention, and multi-dimensional collaborative attention, which effectively extract global channel and temporal-frequency information with low parameters and computational resources. Based on the results from the L3DAS22 dataset, SF-AN achieves a Metric score of 0.933 with only 778.75 K parameters and 4.78 G MACs, outperforming other models that require more computational resources. Under various noise and reverberation conditions, SF-AN also exhibits outstanding denoising and dereverberation performance. On average, the PESQ of SF-AN improves from 3.359 to 3.420 in noisy environments, and from 2.931 to 3.328 in reverberant environments compared to EaBNet. When both noise and reverberation are present, SF-AN achieves a PESQ score of 2.864, a STOI score of 0.887, and a SI-SDR score of 6.233. SF-AN is capable of effectively suppressing noise and reverberation while maintaining a very lightweight structure, thereby enhancing the listening experience for listener.
| Original language | English |
|---|---|
| Article number | 103296 |
| Journal | Speech Communication |
| Volume | 174 |
| DOIs | |
| State | Published - Oct 2025 |
| Externally published | Yes |
Keywords
- Channel shuffling
- Fourier convolution
- Lightweight model
- Multi-channel speech enhancement
Fingerprint
Dive into the research topics of 'SF-AN: A lightweight shuffle Fourier attention network for multi-channel speech enhancement'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver