Skip to main navigation Skip to search Skip to main content

Supervised Attention Multi-Scale Temporal Convolutional Network for monaural speech enhancement

  • Zehua Zhang
  • , Lu Zhang
  • , Xuyi Zhuang
  • , Yukun Qian
  • , Mingjiang Wang*
  • *Corresponding author for this work
  • Harbin Institute of Technology Shenzhen
  • Ltd

Research output: Contribution to journalArticlepeer-review

Abstract

Speech signals are often distorted by reverberation and noise, with a widely distributed signal-to-noise ratio (SNR). To address this, our study develops robust, deep neural network (DNN)-based speech enhancement methods. We reproduce several DNN-based monaural speech enhancement methods and outline a strategy for constructing datasets. This strategy, validated through experimental reproductions, has effectively enhanced the denoising efficiency and robustness of the models. Then, we propose a causal speech enhancement system named Supervised Attention Multi-Scale Temporal Convolutional Network (SA-MSTCN). SA-MSTCN extracts the complex compressed spectrum (CCS) for input encoding and employs complex ratio masking (CRM) for output decoding. The supervised attention module, a lightweight addition to SA-MSTCN, guides feature extraction. Experiment results show that the supervised attention module effectively improves noise reduction performance with a minor increase in computational cost. The multi-scale temporal convolutional network refines the perceptual field and better reconstructs the speech signal. Overall, SA-MSTCN not only achieves state-of-the-art speech quality and intelligibility compared to other methods but also maintains stable denoising performance across various environments.

Original languageEnglish
Article number20
JournalEurasip Journal on Audio, Speech, and Music Processing
Volume2024
Issue number1
DOIs
StatePublished - Dec 2024
Externally publishedYes

Keywords

  • Complex compressed spectrum
  • Complex ratio mask
  • Monaural speech enhancement
  • Multi-scale temporal convolutional network
  • Supervised attention

Fingerprint

Dive into the research topics of 'Supervised Attention Multi-Scale Temporal Convolutional Network for monaural speech enhancement'. Together they form a unique fingerprint.

Cite this