Skip to main navigation Skip to search Skip to main content

RAW-CLIP Fusion: Unleashing Semantic-Aware Denoising for Sensor-Agnostic Low-Light Imaging

  • Mingde Qiao
  • , Junjun Jiang*
  • , Qing Ma
  • , Zhanghong Zhao
  • , Junhui Hou
  • , Jiayi Ma
  • *Corresponding author for this work
  • School of Computer Science and Technology, Harbin Institute of Technology
  • Hong Kong Polytechnic University
  • Zhejiang Chinese Medical University
  • Henan University of Engineering
  • City University of Hong Kong
  • Wuhan University

Research output: Contribution to journalArticlepeer-review

Abstract

Denoising images captured under extreme low-light conditions remains a persistent challenge in computational photography, primarily due to low signal-to-noise ratios and sensor-specific noise characteristics. These variations often require per-sensor noise calibration to achieve effective denoising. Although recent calibration-free methods aim to reduce this dependency through synthetic noise modeling or few-shot fine-tuning, their performance often degrades in extreme low-light scenarios across different sensors due to mismatches between synthetic and real-world noise. To address this gap, we introduce CLIP-Guided Denoising (CLD), the first framework to leverage large-scale vision models pretrained on sRGB images for cross-domain feature fusion, effectively guiding RAW image denoising across diverse sensors. Although not trained on RAW data, CLIP embeddings offer semantically robust and noise-invariant features that help guide the denoising network to focus on the underlying image content rather than fitting to specific noise distributions. Extensive experiments on the SID and ELD datasets demonstrate that CLD achieves state-of-the-art performance in calibration-free settings, significantly outperforming prior methods under extreme low-light conditions and achieving robust generalization across unseen sensor domains.

Original languageEnglish
Pages (from-to)7165-7178
Number of pages14
JournalIEEE Transactions on Image Processing
Volume35
DOIs
StatePublished - 2026
Externally publishedYes

Keywords

  • CLIP
  • Few-shot
  • deep learning
  • low-light RAW image denoising
  • robust denoising model

Fingerprint

Dive into the research topics of 'RAW-CLIP Fusion: Unleashing Semantic-Aware Denoising for Sensor-Agnostic Low-Light Imaging'. Together they form a unique fingerprint.

Cite this