Skip to main navigation Skip to search Skip to main content

M2DAO-Talker: Harmonizing Multi-Granular Motion Decoupling and Alternating Optimization for Talking-Head Generation

  • School of Computer Science and Technology, Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

The generation of high-fidelity, audio-driven talking head videos is critical for immersive digital human applications. While 3D Gaussian Splatting (3DGS) has emerged as a highly efficient alternative to Neural Radiance Fields (NeRF) by offering superior rendering speeds, existing 3DGS-based methods still suffer from visual artifacts - such as temporal jitter, motion discontinuities, and local piercing - primarily caused by imprecise motion decoupling. To overcome these limitations, we propose M2DAO-Talker, a novel framework integrating Multi-granular Motion decoupling and Dynamic Alternating Optimization. Built upon robustly extracted deformation priors, the core of our framework features a multi-granular motion decoupling scheme that explicitly disentangles head, facial, and oral movements into dedicated branches. To ensure seamless integration, we introduce a Composite-level Coherence Supervision (CCS) mechanism that enforces spatial consistency between the dynamic face region and the torso context. Furthermore, a staged Alternating Optimization Strategy (AOS) is proposed to mitigate branch-wise interference and enhance boundary refinement between facial and oral regions. Extensive experiments across multiple datasets demonstrate that M2DAO-Talker significantly outperforms state-of-the-art approaches both qualitatively and quantitatively. Notably, our method achieves a 2.43 dB PSNR improvement in generation quality and a 0.64 increase in user-evaluated realism compared to TalkingGaussian, while maintaining a real-time inference speed of 150 FPS.

Original languageEnglish
Pages (from-to)11817-11830
Number of pages14
JournalIEEE Transactions on Circuits and Systems for Video Technology
Volume36
Issue number8
DOIs
StatePublished - 2026
Externally publishedYes

Keywords

  • Talking-head generation
  • alternating optimization
  • motion decoupling
  • multi-granular representation

Fingerprint

Dive into the research topics of 'M2DAO-Talker: Harmonizing Multi-Granular Motion Decoupling and Alternating Optimization for Talking-Head Generation'. Together they form a unique fingerprint.

Cite this