Abstract
Image manipulation localization (IML) aims to predict pixel-level masks of forged regions in tampered images. Most existing IML approaches are based on deep learning, whose performance strongly depends on large, high-quality datasets. However, the high cost of pixel-level annotation constrains the scale and diversity of current datasets. To overcome data scarcity, constrained IML (CIML) has been introduced to automatically generate masks from original–forgery pairs. Despite different inputs, both IML and CIML target the same objective: estimating the manipulation mask of a forged image. However, existing methods suffer from two limitations: (1) they treat IML and CIML as separate tasks with task-specific architectures despite their shared goal; (2) they are purely discriminative networks that rely on noisy labels and only produce deterministic masks without explicit uncertainty estimation. This study proposes UGD-IML, a generative diffusion framework that models manipulation masks in a continuous embedding space and handles both forged images and original–forgery pairs within a single conditional architecture. Extensive experiments demonstrate state-of-the-art performance, surpassing strong baselines by average F1 gains of 9.66% and 4.36% on IML and CIML, respectively, while also supporting dynamic inference and uncertainty estimation and exhibiting robustness under post-processing operations.
| Original language | English |
|---|---|
| Article number | 134064 |
| Journal | Neurocomputing |
| Volume | 696 |
| DOIs | |
| State | Published - 1 Oct 2026 |
| Externally published | Yes |
Keywords
- Constrained image manipulation localization
- Diffusion model
- Image manipulation localization
Fingerprint
Dive into the research topics of 'UGD-IML: A unified generative diffusion-based framework for constrained and unconstrained image manipulation localization'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver