Skip to main navigation Skip to search Skip to main content

FSTF-AN: Fused Sparse Temporal-Frequency Attentive Network for Multi-Channel Speech Enhancement

  • Shiyun Xu
  • , Yinghan Cao
  • , Wenjie Zhang
  • , Zehua Zhang
  • , Mingjiang Wang*
  • *Corresponding author for this work
  • Harbin Institute of Technology Shenzhen

Research output: Contribution to journalArticlepeer-review

Abstract

The Transformer has achieved impressive performance in the multi-channel speech enhancement field; however, it struggles to capture local features, which leads to the loss of speech details. To enhance the extraction of local features in the network, we propose a fused sparse temporal-frequency attentive network (FSTF-AN), which aims to fully capture features across the temporal-frequency, frequency, and temporal dimensions. We propose top-k fused sparse self-attention, which employs a fusion strategy to adaptively retain the most crucial attention scores when computing self-attention maps, thereby eliminating irrelevant information interference and better aggregating features. Furthermore, we propose a multi-scale fused feed-forward network, which effectively captures multi-scale features, further enhancing the network's ability to capture local features. The experimental results demonstrate that FSTF-AN exhibits significant advantages over other SOTA models, effectively enhancing speech quality and intelligibility.

Original languageEnglish
Pages (from-to)2124-2128
Number of pages5
JournalIEEE Signal Processing Letters
Volume32
DOIs
StatePublished - 2025
Externally publishedYes

Keywords

  • Multi-channel speech enhancement
  • multi-scale feed-forward network
  • sparse self-attention
  • transformer

Fingerprint

Dive into the research topics of 'FSTF-AN: Fused Sparse Temporal-Frequency Attentive Network for Multi-Channel Speech Enhancement'. Together they form a unique fingerprint.

Cite this