Skip to main navigation Skip to search Skip to main content

Text-independent Speaker Recognition Based on X-vector

  • Lianyu Zhou*
  • , Mingjiang Wang
  • , Yukun Qian
  • , Huaiwen Luo
  • , Heng Li
  • , Xu Lin
  • *Corresponding author for this work
  • Harbin Institute of Technology Shenzhen

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Speaker recognition is also called voiceprint recognition. The current state-of-the-art technology for speaker recognition is to use deep neural networks to extract features of the speaker's speech. This embedded feature extracted by DNN is generally called x-vector. Recently, resnet-based structures have received extensive attention and have gradually become the basis for speaker recognition research. In terms of model input, the most commonly used features include Linear Prediction Coefficient, Mel Frequency Cepstral Coefficient, Mel Filter Bank, and Spectrogram. However, a single feature cannot reveal all the features of speech. In this paper, we propose a text-independent speaker recognition algorithm based on fused features and x-vector architecture, in which we use LPC, F-bank and Spectrogram for acoustic features and fuse them at frame level, we use the currently popular ResNet as model for training and modify its structure, we use the additive angular margin loss for classification loss function. The experiments show that our proposed fusion feature and modified ResNet achieves remarkable Equal Error Rate of 0.9 for the VTCK dataset, which greatly improves the accuracy of speaker recognition.

Original languageEnglish
Title of host publication2022 7th International Conference on Signal and Image Processing, ICSIP 2022
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages121-125
Number of pages5
ISBN (Electronic)9781665495639
DOIs
StatePublished - 2022
Externally publishedYes
Event7th International Conference on Signal and Image Processing, ICSIP 2022 - Suzhou, China
Duration: 20 Jul 202222 Jul 2022

Publication series

Name2022 7th International Conference on Signal and Image Processing, ICSIP 2022

Conference

Conference7th International Conference on Signal and Image Processing, ICSIP 2022
Country/TerritoryChina
CitySuzhou
Period20/07/2222/07/22

Keywords

  • F-bank
  • LPC
  • ResNet
  • additive angular margin loss
  • speaker recognition
  • spectrogram
  • text-independent
  • x-vector

Fingerprint

Dive into the research topics of 'Text-independent Speaker Recognition Based on X-vector'. Together they form a unique fingerprint.

Cite this