Skip to main navigation Skip to search Skip to main content

SMACNet: A Unified Framework for One-Shot Talking Head Synthesis via Subtle Motion and Appearance Compensation

  • Yuzhu Ji
  • , Wei Hu
  • , Guoji Gan
  • , Zhiquan Long
  • , Yiqun Zhang*
  • , An Zeng*
  • , Haijun Zhang
  • *Corresponding author for this work
  • Guangdong University of Technology
  • Harbin Institute of Technology Shenzhen

Research output: Contribution to journalArticlepeer-review

Abstract

One-shot talking head synthesis refers to animating a source person’s portrait image using driving video sequences while maintaining the driving pose’s accuracy and preserving the source’s appearance. Although recent facial keypoint-based methods have made significant strides in high-fidelity animations, achieving high-quality cross-identity animation, it remains challenging for subtle facial motion and expression transfer with correct geometry and appearance. To address these limitations, we propose SMACNet, a unified framework designed for subtle motion transfer and fine-grained appearance recovery. In particular, we leverage 3DMM parameters as robust 3D geometry cues for both subtle motion compensation and detailed appearance feature learning. To accomplish this, we introduce a dual-branch subtle motion compensation (DSMC) network to capture subtle facial motions and compensate for overall head pose. Furthermore, we design a pose-constrained appearance feature compensation (PAFC) module to restore detailed facial appearance by modulating features from a learnable appearance feature memory bank. Additionally, a 3DMM coefficient-constrained (CC) module is integrated to preserve appearance consistency by conditioning the synthesis on the source identity. Experimental results show that SMACNet outperforms state-of-the-art methods, generalizes well across datasets and identities, and produces high-quality photo-realistic facial animations with accurate subtle motion transfer and consistent identity preservation.

Original languageEnglish
Pages (from-to)9315-9327
Number of pages13
JournalIEEE Transactions on Consumer Electronics
Volume71
Issue number4
DOIs
StatePublished - 2025
Externally publishedYes

Keywords

  • Neural talking head
  • motion transfer
  • subtle motion compensation
  • video synthesis

Fingerprint

Dive into the research topics of 'SMACNet: A Unified Framework for One-Shot Talking Head Synthesis via Subtle Motion and Appearance Compensation'. Together they form a unique fingerprint.

Cite this