Abstract
In this paper, we propose a cluster adaptive training (CAT) algorithm using sparse training data for textindependent speaker verification. CAT is initiated by a Gaussian Mixture Model-Universal Background Model (GMM-UBM) speaker verification system, followed by a hierarchical clustering. The similarity measure used in the hierarchical clustering is the square root of the weighted sum of the Bhattacharyya distance or symmetric Kullback-Leibler divergence. First, a GMM-UBM system is constructed as a baseline system and used to derive speaker models. Then, CAT is applied to find groups of speakers with similar sounding voices and whose models are highly similar. Finally, each speaker model is re-estimated using both his/her own training data and sound-alike speakers' training data. Using the proposed method, about a relative reduction of 6.3% in the equal error rate is achieved compared with the GMM-UBM baseline system.
| Original language | English |
|---|---|
| Pages (from-to) | 235-239 |
| Number of pages | 5 |
| Journal | Advanced Science Letters |
| Volume | 5 |
| Issue number | 1 |
| DOIs | |
| State | Published - Jan 2012 |
| Externally published | Yes |
Keywords
- Cluster adaptive training
- Gaussian mixture model
- Similarity measure
- Sparse training data
- Speaker verification
Fingerprint
Dive into the research topics of 'A cluster adaptive training algorithm for text-independent speaker verification with sparse training data'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver