Skip to main navigation Skip to search Skip to main content

Wireless image transmission via joint source-channel coding guided by diffusion models and multimodal semantics

  • Fuqiang Liu
  • , Lin Ma
  • , Yun Jia*
  • *Corresponding author for this work
  • Shandong Technology and Business University
  • School of Electronics and Information Engineering, Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Wireless image transmission based on deep joint source-channel coding (Deep JSCC) often suffers from perceptual quality degradation under low signal-to-noise ratio (SNR) conditions. To address this, this paper presents SemDiff-JSCC, a framework guided by multimodal semantics. By leveraging text descriptions and edge maps to guide the diffusion-based denoising process, the proposed approach enhances image reconstruction accuracy over noisy wireless channels. An OmniControl-based parameter reuse strategy is adopted to reduce complexity, alongside a blind channel estimation mechanism for robust adaptation to slow fading channels without dedicated pilots. Experimental results demonstrate that SemDiff-JSCC consistently outperforms baselines, particularly in scenarios with unknown SNR or channel state information. Notably, at 0 dB SNR and an extreme compression ratio of 1/128, it achieves over 21% improvement in the Fréchet inception distance (FID) metric compared with the strongest baseline at the same rate. Furthermore, it attains superior perceptual quality to existing schemes operating at a 8× higher bandwidth (R=1/16), highlighting its efficacy for high-fidelity wireless semantic communication.

Original languageEnglish
Article number033019
JournalJournal of Electronic Imaging
Volume35
Issue number3
DOIs
StatePublished - 1 May 2026
Externally publishedYes

Keywords

  • diffusion models
  • joint source-channel coding
  • semantic communication
  • semantic guidance
  • wireless image transmission

Fingerprint

Dive into the research topics of 'Wireless image transmission via joint source-channel coding guided by diffusion models and multimodal semantics'. Together they form a unique fingerprint.

Cite this