Skip to main navigation Skip to search Skip to main content

OACodec: Audio attribute disentanglement via orthogonal disentanglement and mutual information minimization

  • Yukun Qian
  • , Wenjie Zhang
  • , Zehua Zhang
  • , Lianyu Zhou
  • , Xuyi Zhuang
  • , Mingjiang Wang*
  • *Corresponding author for this work
  • Harbin Institute of Technology Shenzhen

Research output: Contribution to journalArticlepeer-review

Abstract

Neural audio codecs (NACs) based on end-to-end neural networks and vector quantization have recently achieved high-fidelity speech compression and reconstruction. However, most existing NACs learn entangled latent codes that mix speaker timbre, prosody, and phonetic information, which limits interpretability and controllability. Although several attempts introduce attribute-aware objectives, they often lack an explicit decomposition mechanism and a principled information-theoretic constraint to encourage independent factorization. We propose OACodec, a novel neural audio disentanglement codec specifically designed to learn disentangled representations of speech attributes. Unlike conventional neural audio codec systems that treat audio as undifferentiated data, OACodec introduces a multi-stage orthogonal disentanglement network that explicitly separates timbre, prosody and phonetic information. Each latent attribute is extracted via a residual separation mechanism, guided by orthogonality constraints and supervised learning. To further promote independent factorization, we employ mutual information estimators both within and across attribute components, minimizing their mutual information. Additionally, we adopt a smoothed Tchebycheff optimization strategy to achieve a Pareto-optimal balance between reconstruction fidelity and disentanglement objectives. Experimental results demonstrate the effectiveness of OACodec in both zero-shot voice conversion and speech reconstruction, where it outperforms VC baselines and FACodec in zero-shot VC tasks and achieves reconstruction quality comparable to EnCodec and HiFi-Codec while surpassing FACodec. We further demonstrate, by constructing a two-stage TTS model, that the disentangled codes effectively improve performance on the TTS task. Ablation studies further validate the contribution of each proposed module. OACodec lays a strong foundation for interpretable and controllable speech modeling, with promising implications for applications such as text-to-speech and speech-based language modeling. Audio samples are available at https://hamidun123.github.io/OACodecDemo/.

Original languageEnglish
Article number109128
JournalNeural Networks
Volume203
DOIs
StatePublished - Nov 2026
Externally publishedYes

Keywords

  • Disentangled representation learning
  • Mutual information
  • Neural audio codec
  • Orthogonal disentanglement

Fingerprint

Dive into the research topics of 'OACodec: Audio attribute disentanglement via orthogonal disentanglement and mutual information minimization'. Together they form a unique fingerprint.

Cite this