Skip to main navigation Skip to search Skip to main content

CARVE: Content-Adaptive Rate-Variable Encoding for Neural Speech Codecs

  • Wenjie Zhang
  • , Yukun Qian
  • , Yinghan Cao
  • , Changjun He
  • , Shiyun Xu
  • , Mingjiang Wang*
  • *Corresponding author for this work
  • Harbin Institute of Technology Shenzhen

Research output: Contribution to journalArticlepeer-review

Abstract

Neural speech codecs that convert continuous waveforms into discrete tokens are an essential component for building compact speech representations in downstream applications. However, most existing codecs require both fixed and high frame rates to maintain reconstruction quality, which reduces the efficiency of downstream applications. To address this issue, we propose CARVE, a Content-Adaptive Rate-Variable Encoding strategy that performs compression at an externally specified frame rate within the latent space of a neural speech codec. CARVE comprises three modules: (i) Unsupervised Density-Aware Scoring Module that combines inter-frame similarity with redundancy-aware cues to quantify the relative importance of each frame; (ii) Sparse Frame Segmentation Module that allocates finer segmentation to high-density information regions and coarser segmentation to low-density information regions based on these scores; and (iii) Context-Aware Frame Reconstruction Module that maps segment-level context back to frame-level latent representations. Integrating CARVE into the neural speech codec, extensive experiments show that it can perform compression at an externally specified frame rate while maintaining high reconstruction quality, thereby validating its effectiveness and practicality in low-frame-rate neural speech coding scenarios.

Original languageEnglish
Pages (from-to)2036-2040
Number of pages5
JournalIEEE Signal Processing Letters
Volume33
DOIs
StatePublished - 2026
Externally publishedYes

Keywords

  • Neural speech codec
  • information density
  • speech coding
  • variable frame rate

Fingerprint

Dive into the research topics of 'CARVE: Content-Adaptive Rate-Variable Encoding for Neural Speech Codecs'. Together they form a unique fingerprint.

Cite this