Skip to main navigation Skip to search Skip to main content

Zero-shot cross-domain image composition via self attention injection

  • Xiangrui Chen
  • , Qi Si
  • , Bo Wang
  • , Zhao Zhang*
  • , Xianming Ye
  • , Yun Yang
  • , Haijun Zhang
  • *Corresponding author for this work
  • Hefei University of Technology
  • Hebei University of Technology
  • Yunnan Key Laboratory of Software Engineering
  • University of Pretoria
  • Yunnan University
  • Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Leveraging the robust generative priors of diffusion models, image composition has achieved remarkable progress. However, existing approaches continue to grapple with a persistent dilemma: the trade-off between maintaining the structural fidelity of the source object and achieving deep stylistic harmonization with the background. We attribute this limitation to two primary factors: 1) the insufficient disentanglement of geometric structure and visual appearance in current architectures, leading to conflicts during the generation process; and 2) the reliance on global statistical alignment techniques, which merely adjust tonal distributions but fail to capture complex semantic stylistic patterns. To address these challenges, we propose a novel training-free tri-branch denoising framework that effectively decouples structure from style via attention manipulation. Specifically, we propose two core mechanisms. Semantic Injection employs self attention maps to separate an object’s spatial structure from its visual appearance. Style Guidance adapts advanced attention based style transfer techniques to the composition task for the first time. Comprehensive experimental results show that our method outperforms existing state-of-the-art approaches and achieves consistent improvements in structural consistency and stylistic coherence for image composition.

Original languageEnglish
Article number109412
JournalNeural Networks
Volume205
DOIs
StatePublished - Jan 2027
Externally publishedYes

Keywords

  • Diffusion models
  • Image composition
  • Tuning-free

Fingerprint

Dive into the research topics of 'Zero-shot cross-domain image composition via self attention injection'. Together they form a unique fingerprint.

Cite this