Abstract
The Transformer has achieved impressive performance in the multi-channel speech enhancement field; however, it struggles to capture local features, which leads to the loss of speech details. To enhance the extraction of local features in the network, we propose a fused sparse temporal-frequency attentive network (FSTF-AN), which aims to fully capture features across the temporal-frequency, frequency, and temporal dimensions. We propose top-k fused sparse self-attention, which employs a fusion strategy to adaptively retain the most crucial attention scores when computing self-attention maps, thereby eliminating irrelevant information interference and better aggregating features. Furthermore, we propose a multi-scale fused feed-forward network, which effectively captures multi-scale features, further enhancing the network's ability to capture local features. The experimental results demonstrate that FSTF-AN exhibits significant advantages over other SOTA models, effectively enhancing speech quality and intelligibility.
| Original language | English |
|---|---|
| Pages (from-to) | 2124-2128 |
| Number of pages | 5 |
| Journal | IEEE Signal Processing Letters |
| Volume | 32 |
| DOIs | |
| State | Published - 2025 |
| Externally published | Yes |
Keywords
- Multi-channel speech enhancement
- multi-scale feed-forward network
- sparse self-attention
- transformer
Fingerprint
Dive into the research topics of 'FSTF-AN: Fused Sparse Temporal-Frequency Attentive Network for Multi-Channel Speech Enhancement'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver