Skip to main navigation Skip to search Skip to main content

Real-world speech/non-speech audio classification based on sparse representation features and GPCs

  • School of Computer Science and Technology, Harbin Institute of Technology

Research output: Contribution to journalConference articlepeer-review

Abstract

A novel and robust approach for content based speech/nonspeech audio classification is proposed based on sparse representation (SR) features and Gaussian process classifiers (GPCs). The projections of the noise robust sparse representations for audio signals computed by L1 -norm minimization are used as features. GPCs are used to learn and predict audio categories. Compare to the difficulties of Support Vector Machines (SVMs) in determining the hyperparameters, GPCs employ Bayesian selection criterion to estimate them. Experimental results on real-world audio datasets show that the SR features are more robust to audio variants than mel-frequency cepstral coefficients (MFCCs) and the proposed approach gives better performances than SVM.

Original languageEnglish
Pages (from-to)2401-2404
Number of pages4
JournalProceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
DOIs
StatePublished - 2011
Externally publishedYes
Event12th Annual Conference of the International Speech Communication Association, INTERSPEECH 2011 - Florence, Italy
Duration: 27 Aug 201131 Aug 2011

Keywords

  • Audio classification
  • Gaussian process classifiers
  • Sparse representation
  • Speech discrimination

Fingerprint

Dive into the research topics of 'Real-world speech/non-speech audio classification based on sparse representation features and GPCs'. Together they form a unique fingerprint.

Cite this