Skip to main navigation Skip to search Skip to main content

DiTalker: A unified DiT-based framework for high-quality and style-controllable portrait animation

  • He Feng
  • , Yongjia Ma
  • , Lei Fan
  • , Donglin Di
  • , Tonghua Su*
  • *Corresponding author for this work
  • Harbin Institute of Technology
  • Li Auto Inc.
  • University of New South Wales

Research output: Contribution to journalArticlepeer-review

Abstract

Portrait animation aims to synthesize talking videos from a static reference face, conditioned on audio and style frames, while ensuring precise lip synchronization and faithful reproduction of speaking styles. Existing diffusion-based portrait animation methods primarily focus on lip synchronization or static emotion transformation, often overlooking dynamic styles such as head movements. Moreover, most of these methods rely on a dual U-Net architecture, which preserves identity consistency but incurs additional architectural complexity. To this end, we propose DiTalker, a unified DiT-based framework for style-controllable portrait animation. We design a Style-Expression Encoding Module that employs two separate branches: a style branch extracting identity-specific facial dynamics, and an expression branch extracting global expression semantics. We further introduce an Audio-Style Fusion Module that processes audio and style features in parallel via two cross-attention layers, whose outputs are scaled and fused through element-wise addition to guide the animation process. To ensure the quality of animation results, we adopt and modify two optimization constraints: one to improve lip synchronization and the other to preserve fine-grained identity and background details. Extensive experiments demonstrate the superiority of DiTalker in terms of lip synchronization and speaking style controllability. Code is available at https://github.com/fenghe12/DiTalker-CVIU/ .

Original languageEnglish
Article number104819
JournalComputer Vision and Image Understanding
Volume269
DOIs
StatePublished - Jun 2026

Keywords

  • Diffusion transformer
  • Lip synchronization
  • Portrait animation
  • Speaking style controllable animation
  • Video generation

Fingerprint

Dive into the research topics of 'DiTalker: A unified DiT-based framework for high-quality and style-controllable portrait animation'. Together they form a unique fingerprint.

Cite this