Abstract
Speech recognition systems obtaining high recognition rates in clean environments perform badly in mismatch environments without compensation. Based on the research, we found that bandwidth mismatch, namely the bandwidth difference between the training and test conditions, is one of the main factors leading to environment mismatch. When the bandwidth of the test speech is narrower than that of the training speech, the distortion is non-invertible and time-varying in the logarithm spectrum and cepstrum domains. So it could not be compensated with current channel compensation methods. After analyzing the Mel-frequency cepstrum coefficient distortion caused by the lost frequency band, we propose a compensation method based on spectral fold. Furthermore, we provide an algorithm for speech bandwidth detection and a unified compensation framework. Experiments on the AN4 and TIMIT/TIMIT databases show that the proposed framework improved the robustness of speech recognition underbandwidth mismatch conditions.
| Original language | English |
|---|---|
| Pages (from-to) | 1629-1637 |
| Number of pages | 9 |
| Journal | Jisuanji Xuebao/Chinese Journal of Computers |
| Volume | 34 |
| Issue number | 9 |
| DOIs | |
| State | Published - Sep 2011 |
| Externally published | Yes |
Keywords
- Bandwidth mismatch
- Distortion compensation
- MFCC
- Robustness
- Speech recognition
Fingerprint
Dive into the research topics of 'Research on bandwidth mismatch compensation in speech recognition'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver