Abstract
Fine-grained object-centric editing in complex scenes, while preserving contextual integrity, remains a persistent challenge. The core difficulties arise from two sources: (1) inaccurate object localization stemming from cross-attention misalignment and inter-object interference in diffusion models, where imperfect attention correspondence frequently yields incomplete or misplaced edits; and (2) the reliance of mask-conditioned diffusion on random Gaussian noise for generating content within edited regions, which affords limited control over precise object placement. To tackle these issues, this paper proposes a training-free framework grounded in latent diffusion models. Concretely, we introduce a latent space optimization strategy that refines cross-attention maps to disentangle object representations and achieve accurate spatial alignment, dynamically adjusting attention weights across distinct objects to suppress mutual interference. Furthermore, we design a region-aware fusion mechanism to safeguard background structure and content during editing, adaptively blending the edited latent features with the original background information to prevent structural distortion. Experimental evaluations on public benchmarks demonstrate that the proposed method consistently outperforms state-of-the-art approaches, delivering clear gains in both structural fidelity and semantic coherence.
| Original language | English |
|---|---|
| Article number | 109058 |
| Journal | Neural Networks |
| Volume | 202 |
| DOIs | |
| State | Published - Oct 2026 |
| Externally published | Yes |
Keywords
- Attention guidance
- Diffusion model
- Energy-guided latent optimization
- Object-centric editing
- Training-free
Fingerprint
Dive into the research topics of 'Object-centric image editing via position-structure guided diffusion'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver