TY - GEN
T1 - Text-independent Speaker Recognition Based on X-vector
AU - Zhou, Lianyu
AU - Wang, Mingjiang
AU - Qian, Yukun
AU - Luo, Huaiwen
AU - Li, Heng
AU - Lin, Xu
N1 - Publisher Copyright:
© 2022 IEEE.
PY - 2022
Y1 - 2022
N2 - Speaker recognition is also called voiceprint recognition. The current state-of-the-art technology for speaker recognition is to use deep neural networks to extract features of the speaker's speech. This embedded feature extracted by DNN is generally called x-vector. Recently, resnet-based structures have received extensive attention and have gradually become the basis for speaker recognition research. In terms of model input, the most commonly used features include Linear Prediction Coefficient, Mel Frequency Cepstral Coefficient, Mel Filter Bank, and Spectrogram. However, a single feature cannot reveal all the features of speech. In this paper, we propose a text-independent speaker recognition algorithm based on fused features and x-vector architecture, in which we use LPC, F-bank and Spectrogram for acoustic features and fuse them at frame level, we use the currently popular ResNet as model for training and modify its structure, we use the additive angular margin loss for classification loss function. The experiments show that our proposed fusion feature and modified ResNet achieves remarkable Equal Error Rate of 0.9 for the VTCK dataset, which greatly improves the accuracy of speaker recognition.
AB - Speaker recognition is also called voiceprint recognition. The current state-of-the-art technology for speaker recognition is to use deep neural networks to extract features of the speaker's speech. This embedded feature extracted by DNN is generally called x-vector. Recently, resnet-based structures have received extensive attention and have gradually become the basis for speaker recognition research. In terms of model input, the most commonly used features include Linear Prediction Coefficient, Mel Frequency Cepstral Coefficient, Mel Filter Bank, and Spectrogram. However, a single feature cannot reveal all the features of speech. In this paper, we propose a text-independent speaker recognition algorithm based on fused features and x-vector architecture, in which we use LPC, F-bank and Spectrogram for acoustic features and fuse them at frame level, we use the currently popular ResNet as model for training and modify its structure, we use the additive angular margin loss for classification loss function. The experiments show that our proposed fusion feature and modified ResNet achieves remarkable Equal Error Rate of 0.9 for the VTCK dataset, which greatly improves the accuracy of speaker recognition.
KW - F-bank
KW - LPC
KW - ResNet
KW - additive angular margin loss
KW - speaker recognition
KW - spectrogram
KW - text-independent
KW - x-vector
UR - https://www.scopus.com/pages/publications/85139397673
U2 - 10.1109/ICSIP55141.2022.9887021
DO - 10.1109/ICSIP55141.2022.9887021
M3 - 会议稿件
AN - SCOPUS:85139397673
T3 - 2022 7th International Conference on Signal and Image Processing, ICSIP 2022
SP - 121
EP - 125
BT - 2022 7th International Conference on Signal and Image Processing, ICSIP 2022
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 7th International Conference on Signal and Image Processing, ICSIP 2022
Y2 - 20 July 2022 through 22 July 2022
ER -