Skip to main navigation Skip to search Skip to main content

Cassic: Towards Content-Adaptive State-Space Models for Learned Image Compression

  • Shiyu Qin
  • , Jinpeng Wang
  • , Yimin Zhou
  • , Bin Chen*
  • , Tianci Luo
  • , Baoyi An
  • , Tao Dai
  • , Shu Tao Xia
  • , Yaowei Wang
  • *Corresponding author for this work
  • Tsinghua University
  • Harbin Institute of Technology Shenzhen
  • Peng Cheng Laboratory
  • Huawei Technologies Co., Ltd.
  • Shenzhen University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Learned image compression (LIC) demonstrates superior rate-distortion (RD) performance compared to traditional methods. Recent method MambaVC attempts to introduce Mamba, a variant of state space models, into this field aim to establish a new paradigm beyond convolutional neural networks and transformers. However, this approach relies on predefined four-directional scanning, which prioritizes spatial proximity over content and semantic relationships, resulting in suboptimal redundancy elimination. Additionally, it focuses solely on nonlinear transformations, neglecting entropy model improvements crucial for accurate probability estimation in entropy coding. To address these limitations, we propose a novel framework based on content-adaptive visual state space model, Cassic, through dual innovation. First, we design a content-adaptive selective scan based on weighted activation maps and bit allocation maps, subsequently developing a content-adaptive visual state space block. Second, we present a mamba-based channel-wise auto-regressive entropy model to fully leverage inter-slice bit allocation consistency for enhanced probability estimation. Extensive experimental results demonstrate that our method achieves state-of-the-art performance across three datasets while maintaining faster processing speeds than existing MambaVC approach.

Original languageEnglish
Title of host publicationProceedings - 2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages15727-15736
Number of pages10
ISBN (Electronic)9798331587758
DOIs
StatePublished - 2025
Externally publishedYes
Event2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025 - Honolulu, United States
Duration: 19 Oct 202523 Oct 2025

Publication series

NameProceedings of the IEEE International Conference on Computer Vision
ISSN (Print)1550-5499
ISSN (Electronic)2380-7504

Conference

Conference2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
Country/TerritoryUnited States
CityHonolulu
Period19/10/2523/10/25

Keywords

  • learned image compression
  • state-space models

Fingerprint

Dive into the research topics of 'Cassic: Towards Content-Adaptive State-Space Models for Learned Image Compression'. Together they form a unique fingerprint.

Cite this