Abstract
As the dominant spatial audio format for headphones, binaural audio can convey immersive spatial cues, including source direction, distance, and motion, and is widely used in games, virtual reality, and other immersive scenarios. However, high-quality binaural audio is still mainly obtained via specialized recording setups or by spatializing pre-recorded audio, which limits the flexibility and diversity of binaural audio creation. Recently, text-to-audio generation has shown tremendous potential and promise in content generation, but existing methods mainly focus on mono or simple stereo signals, in which they rarely utilize the specific structure of binaural audio and usually ignore the spatial attributes. To address this limitation, we propose MS-TTBA, a neural mid–side tokenized text-to-binaural audio framework with large language models. Specifically, binaural audio is first converted into mid and side channels and then mapped to highly compressed discrete tokens using the proposed M–S tokenizer, with mid tokens mainly encoding audio content and side tokens emphasizing interaural differences. A Qwen2.5-based decoder-only large language model (LLM) is employed to autoregressively model the joint distribution of mid-side tokens conditioned on text. At this stage, the LLM mainly models long-term semantics and coarse spatial structure. We further design a side-aware gating flow matching (SAGFM) module to refine acoustic and spatial details in the continuous domain. By focusing on segments with salient binaural differences, SAGFM helps generate binaural audio with clearer spatial cues. Experimental results on the text–audio subset of BEWO-1M show that the proposed method outperforms existing text-to-binaural audio baselines on both multiple objective metrics and subjective listening tests, generating binaural audio with more consistent localization and comparable overall perceptual quality.
| Original language | English |
|---|---|
| Article number | 134172 |
| Journal | Neurocomputing |
| Volume | 696 |
| DOIs | |
| State | Published - 1 Oct 2026 |
Keywords
- Binaural audio generation
- Flow matching
- LLM
Fingerprint
Dive into the research topics of 'MS-TTBA: Mid–side tokenized text-to-binaural audio with large language models'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver