Abstract
Denoising images captured under extreme low-light conditions remains a persistent challenge in computational photography, primarily due to low signal-to-noise ratios and sensor-specific noise characteristics. These variations often require per-sensor noise calibration to achieve effective denoising. Although recent calibration-free methods aim to reduce this dependency through synthetic noise modeling or few-shot fine-tuning, their performance often degrades in extreme low-light scenarios across different sensors due to mismatches between synthetic and real-world noise. To address this gap, we introduce CLIP-Guided Denoising (CLD), the first framework to leverage large-scale vision models pretrained on sRGB images for cross-domain feature fusion, effectively guiding RAW image denoising across diverse sensors. Although not trained on RAW data, CLIP embeddings offer semantically robust and noise-invariant features that help guide the denoising network to focus on the underlying image content rather than fitting to specific noise distributions. Extensive experiments on the SID and ELD datasets demonstrate that CLD achieves state-of-the-art performance in calibration-free settings, significantly outperforming prior methods under extreme low-light conditions and achieving robust generalization across unseen sensor domains.
| Original language | English |
|---|---|
| Pages (from-to) | 7165-7178 |
| Number of pages | 14 |
| Journal | IEEE Transactions on Image Processing |
| Volume | 35 |
| DOIs | |
| State | Published - 2026 |
| Externally published | Yes |
Keywords
- CLIP
- Few-shot
- deep learning
- low-light RAW image denoising
- robust denoising model
Fingerprint
Dive into the research topics of 'RAW-CLIP Fusion: Unleashing Semantic-Aware Denoising for Sensor-Agnostic Low-Light Imaging'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver