Abstract
Neural speech codecs that convert continuous waveforms into discrete tokens are an essential component for building compact speech representations in downstream applications. However, most existing codecs require both fixed and high frame rates to maintain reconstruction quality, which reduces the efficiency of downstream applications. To address this issue, we propose CARVE, a Content-Adaptive Rate-Variable Encoding strategy that performs compression at an externally specified frame rate within the latent space of a neural speech codec. CARVE comprises three modules: (i) Unsupervised Density-Aware Scoring Module that combines inter-frame similarity with redundancy-aware cues to quantify the relative importance of each frame; (ii) Sparse Frame Segmentation Module that allocates finer segmentation to high-density information regions and coarser segmentation to low-density information regions based on these scores; and (iii) Context-Aware Frame Reconstruction Module that maps segment-level context back to frame-level latent representations. Integrating CARVE into the neural speech codec, extensive experiments show that it can perform compression at an externally specified frame rate while maintaining high reconstruction quality, thereby validating its effectiveness and practicality in low-frame-rate neural speech coding scenarios.
| Original language | English |
|---|---|
| Pages (from-to) | 2036-2040 |
| Number of pages | 5 |
| Journal | IEEE Signal Processing Letters |
| Volume | 33 |
| DOIs | |
| State | Published - 2026 |
| Externally published | Yes |
Keywords
- Neural speech codec
- information density
- speech coding
- variable frame rate
Fingerprint
Dive into the research topics of 'CARVE: Content-Adaptive Rate-Variable Encoding for Neural Speech Codecs'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver