TY - GEN
T1 - Furcax
T2 - 44th IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2019
AU - Shi, Ziqiang
AU - Lin, Huibin
AU - Liu, Liu
AU - Liu, Rujie
AU - Hayakawa, Shoji
AU - Han, Jiqing
N1 - Publisher Copyright:
© 2019 IEEE.
PY - 2019/5
Y1 - 2019/5
N2 - Deep gated convolutional networks have been proved to be very effective in single channel speech separation. However current state-of-the-art framework often considers training the gated convolutional networks in time-frequency (TF) domain. Such an approach will result in limited perceptual score, such as signal-to-distortion ratio (SDR) upper bound of separated utterances and also fail to exploit an end-to-end framework. In this paper we present an integrated simple and effective end-to-end approach called FurcaX1 to monaural speech separation, which consists of deep gated (de)convolutional neural networks (GCNN) that takes the mixed utterance of two speakers and maps it to two separated utterances, where each utterance contains only one speaker's voice. For the objective, we propose to train the network by directly optimizing utterance level SDR in a permutation invariant training (PIT) style. We execute generative adversarial training (GAT) throughout the training, which makes the separated speech indistinguishable from the real one. Our experiments on the the public WSJ0-2mix data corpus demonstrate that this new scheme can produce more discriminative separated utterances and leading to performance improvement on the speaker separation task.
AB - Deep gated convolutional networks have been proved to be very effective in single channel speech separation. However current state-of-the-art framework often considers training the gated convolutional networks in time-frequency (TF) domain. Such an approach will result in limited perceptual score, such as signal-to-distortion ratio (SDR) upper bound of separated utterances and also fail to exploit an end-to-end framework. In this paper we present an integrated simple and effective end-to-end approach called FurcaX1 to monaural speech separation, which consists of deep gated (de)convolutional neural networks (GCNN) that takes the mixed utterance of two speakers and maps it to two separated utterances, where each utterance contains only one speaker's voice. For the objective, we propose to train the network by directly optimizing utterance level SDR in a permutation invariant training (PIT) style. We execute generative adversarial training (GAT) throughout the training, which makes the separated speech indistinguishable from the real one. Our experiments on the the public WSJ0-2mix data corpus demonstrate that this new scheme can produce more discriminative separated utterances and leading to performance improvement on the speaker separation task.
KW - Speech separation
KW - cocktail party problem
KW - gated convolutional neural network
KW - generative adversarial training
KW - permutation invariant training
UR - https://www.scopus.com/pages/publications/85068979126
U2 - 10.1109/ICASSP.2019.8682429
DO - 10.1109/ICASSP.2019.8682429
M3 - 会议稿件
AN - SCOPUS:85068979126
T3 - ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings
SP - 6985
EP - 6989
BT - 2019 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2019 - Proceedings
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 12 May 2019 through 17 May 2019
ER -