Skip to main navigation Skip to search Skip to main content

MS-TTBA: Mid–side tokenized text-to-binaural audio with large language models

  • Harbin Institute of Technology
  • Harbin Institute of Technology Shenzhen

Research output: Contribution to journalArticlepeer-review

Abstract

As the dominant spatial audio format for headphones, binaural audio can convey immersive spatial cues, including source direction, distance, and motion, and is widely used in games, virtual reality, and other immersive scenarios. However, high-quality binaural audio is still mainly obtained via specialized recording setups or by spatializing pre-recorded audio, which limits the flexibility and diversity of binaural audio creation. Recently, text-to-audio generation has shown tremendous potential and promise in content generation, but existing methods mainly focus on mono or simple stereo signals, in which they rarely utilize the specific structure of binaural audio and usually ignore the spatial attributes. To address this limitation, we propose MS-TTBA, a neural mid–side tokenized text-to-binaural audio framework with large language models. Specifically, binaural audio is first converted into mid and side channels and then mapped to highly compressed discrete tokens using the proposed M–S tokenizer, with mid tokens mainly encoding audio content and side tokens emphasizing interaural differences. A Qwen2.5-based decoder-only large language model (LLM) is employed to autoregressively model the joint distribution of mid-side tokens conditioned on text. At this stage, the LLM mainly models long-term semantics and coarse spatial structure. We further design a side-aware gating flow matching (SAGFM) module to refine acoustic and spatial details in the continuous domain. By focusing on segments with salient binaural differences, SAGFM helps generate binaural audio with clearer spatial cues. Experimental results on the text–audio subset of BEWO-1M show that the proposed method outperforms existing text-to-binaural audio baselines on both multiple objective metrics and subjective listening tests, generating binaural audio with more consistent localization and comparable overall perceptual quality.

Original languageEnglish
Article number134172
JournalNeurocomputing
Volume696
DOIs
StatePublished - 1 Oct 2026

Keywords

  • Binaural audio generation
  • Flow matching
  • LLM

Fingerprint

Dive into the research topics of 'MS-TTBA: Mid–side tokenized text-to-binaural audio with large language models'. Together they form a unique fingerprint.

Cite this