Skip to main navigation Skip to search Skip to main content

CalibCLIP: Contextual Calibration of Dominant Semantics for Text-Driven Image Retrieval

  • Bin Kang
  • , Bin Chen*
  • , Junjie Wang
  • , Yulin Li
  • , Junzhi Zhao
  • , Junle Wang
  • , Zhuotao Tian*
  • *Corresponding author for this work
  • CAS - Chengdu Institute of Computer Application
  • University of Chinese Academy of Sciences
  • International Research Institute for Artificial Intelligence, Harbin Institute of Technology Shenzhen
  • Harbin Institute of Technology Shenzhen
  • Southwest Jiaotong University
  • Tencent

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Existing Visual Language Models (VLMs) suffer structural limitations where a few low contribution tokens may excessively capture global semantics, dominating the information aggregation process and suppressing the discriminative features in text-driven image retrieval tasks. To address this, we introduce CalibCLIP, a training-free method designed to calibrate the suppressive effects of dominant tokens. Specifically, in the visual space, we propose the Contrastive Visual Enhancer (CVE), which decouples visual features into target and low information regions. Subsequently, it identifies dominant tokens and dynamically suppresses their representations. In the textual space, we introduce the Discriminative Concept Calibrator (DCC), which aims to differentiate between general and discriminative concepts within the text query. By mitigating the challenges posed by generic concepts and improving the representations of discriminative concepts, DCC strengthens the differentiation among similar samples. Finally, extensive experiments demonstrate consistent improvements across seven benchmarks spanning three image retrieval tasks, underscoring the effectiveness of CalibCLIP. Code is available at: https://github.com/kangbin98/CalibCLIP.

Original languageEnglish
Title of host publicationMM 2025 - Proceedings of the 33rd ACM International Conference on Multimedia, Co-Located with MM 2025
PublisherAssociation for Computing Machinery, Inc
Pages5140-5149
Number of pages10
ISBN (Electronic)9798400720352
DOIs
StatePublished - 27 Oct 2025
Externally publishedYes
Event33rd ACM International Conference on Multimedia, MM 2025 - Dublin, Ireland
Duration: 27 Oct 202531 Oct 2025

Publication series

NameMM 2025 - Proceedings of the 33rd ACM International Conference on Multimedia, Co-Located with MM 2025

Conference

Conference33rd ACM International Conference on Multimedia, MM 2025
Country/TerritoryIreland
CityDublin
Period27/10/2531/10/25

Keywords

  • text-driven image retrieval
  • visual language models

Fingerprint

Dive into the research topics of 'CalibCLIP: Contextual Calibration of Dominant Semantics for Text-Driven Image Retrieval'. Together they form a unique fingerprint.

Cite this