Skip to main navigation Skip to search Skip to main content

T2ICount: Enhancing Cross-Modal Understanding for Zero-Shot Counting

  • Yifei Qian
  • , Zhongliang Guo
  • , Bowen Deng
  • , Chun Tong Lei
  • , Shuai Zhao
  • , Chun Pong Lau
  • , Xiaopeng Hong
  • , Michael P. Pound*
  • *Corresponding author for this work
  • University of Nottingham
  • University of St Andrews
  • City University of Hong Kong
  • Nanyang Technological University

Research output: Contribution to journalConference articlepeer-review

Abstract

Zero-Shot object counting aims to count instances of arbitrary object categories specified by text descriptions. Existing methods typically rely on vision-language models like CLIP, but often exhibit limited sensitivity to text prompts. We present T21 Count, a diffusion-based framework that lever-ages rich prior knowledge and fine-grained visual understanding from pretrained diffusion models. While one-step demising ensures efficiency, it leads to weakened text sensitivity. To address this challenge, we propose a Hierarchical Semantic Correction Module that progressively refines text-image feature alignment, and a Representational Regional Coherence Loss that provides reliable supervision signals by leveraging the cross-attention maps extracted from the demising U-Net. Furthermore, we observe that current benchmarks mainly focus on majority objects in images, potentially masking models' text sensitivity. To address this, we contribute a challenging re-annotated subset of FSC147 for better evaluation of text-guided counting ability. Extensive experiments demonstrate that our method achieves superior performance across different benchmarks. Code is available at https://github.com/chal5yq/T2lCount.

Original languageEnglish
Pages (from-to)25336-25345
Number of pages10
JournalProceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition
DOIs
StatePublished - 2025
Event2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2025 - Nashville, United States
Duration: 11 Jun 202515 Jun 2025

Fingerprint

Dive into the research topics of 'T2ICount: Enhancing Cross-Modal Understanding for Zero-Shot Counting'. Together they form a unique fingerprint.

Cite this