Skip to main navigation Skip to search Skip to main content

Object-centric image editing via position-structure guided diffusion

  • Qi Si
  • , Xiangrui Chen
  • , Bo Wang
  • , Zhao Zhang*
  • , Mingbo Zhao
  • , Yun Yang
  • , Haijun Zhang
  • *Corresponding author for this work
  • Hefei University of Technology
  • Hebei University of Technology
  • Hebei Key Laboratory of Big Data Calculation
  • Yunnan Key Laboratory of Software Engineering
  • Donghua University
  • Yunnan University
  • Harbin Institute of Technology Shenzhen

Research output: Contribution to journalArticlepeer-review

Abstract

Fine-grained object-centric editing in complex scenes, while preserving contextual integrity, remains a persistent challenge. The core difficulties arise from two sources: (1) inaccurate object localization stemming from cross-attention misalignment and inter-object interference in diffusion models, where imperfect attention correspondence frequently yields incomplete or misplaced edits; and (2) the reliance of mask-conditioned diffusion on random Gaussian noise for generating content within edited regions, which affords limited control over precise object placement. To tackle these issues, this paper proposes a training-free framework grounded in latent diffusion models. Concretely, we introduce a latent space optimization strategy that refines cross-attention maps to disentangle object representations and achieve accurate spatial alignment, dynamically adjusting attention weights across distinct objects to suppress mutual interference. Furthermore, we design a region-aware fusion mechanism to safeguard background structure and content during editing, adaptively blending the edited latent features with the original background information to prevent structural distortion. Experimental evaluations on public benchmarks demonstrate that the proposed method consistently outperforms state-of-the-art approaches, delivering clear gains in both structural fidelity and semantic coherence.

Original languageEnglish
Article number109058
JournalNeural Networks
Volume202
DOIs
StatePublished - Oct 2026
Externally publishedYes

Keywords

  • Attention guidance
  • Diffusion model
  • Energy-guided latent optimization
  • Object-centric editing
  • Training-free

Fingerprint

Dive into the research topics of 'Object-centric image editing via position-structure guided diffusion'. Together they form a unique fingerprint.

Cite this