Abstract
Portrait animation aims to synthesize talking videos from a static reference face, conditioned on audio and style frames, while ensuring precise lip synchronization and faithful reproduction of speaking styles. Existing diffusion-based portrait animation methods primarily focus on lip synchronization or static emotion transformation, often overlooking dynamic styles such as head movements. Moreover, most of these methods rely on a dual U-Net architecture, which preserves identity consistency but incurs additional architectural complexity. To this end, we propose DiTalker, a unified DiT-based framework for style-controllable portrait animation. We design a Style-Expression Encoding Module that employs two separate branches: a style branch extracting identity-specific facial dynamics, and an expression branch extracting global expression semantics. We further introduce an Audio-Style Fusion Module that processes audio and style features in parallel via two cross-attention layers, whose outputs are scaled and fused through element-wise addition to guide the animation process. To ensure the quality of animation results, we adopt and modify two optimization constraints: one to improve lip synchronization and the other to preserve fine-grained identity and background details. Extensive experiments demonstrate the superiority of DiTalker in terms of lip synchronization and speaking style controllability. Code is available at https://github.com/fenghe12/DiTalker-CVIU/ .
| Original language | English |
|---|---|
| Article number | 104819 |
| Journal | Computer Vision and Image Understanding |
| Volume | 269 |
| DOIs | |
| State | Published - Jun 2026 |
Keywords
- Diffusion transformer
- Lip synchronization
- Portrait animation
- Speaking style controllable animation
- Video generation
Fingerprint
Dive into the research topics of 'DiTalker: A unified DiT-based framework for high-quality and style-controllable portrait animation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver