Skip to main navigation Skip to search Skip to main content

SeqCount: A sequence modeling framework for class-agnostic counting

  • Yiming Zhao
  • , Guorong Li*
  • , Xinyan Liu
  • , Xiang He
  • , Zhenjun Han
  • , Yuankai Qi
  • *Corresponding author for this work
  • University of Chinese Academy of Sciences
  • School of Computer Science and Technology (School of Software), Harbin Institute of Technology Weihai
  • Macquarie University

Research output: Contribution to journalArticlepeer-review

Abstract

Class-agnostic counting (CAC) aims to count the number of objects in any category with only a few exemplars. It is crucial for solving the challenge of counting any visual class without re-finetuning, which in turn lowers deployment costs across diverse scenarios. Existing CAC methods usually formulate counting as density map regression, where the supervisory signal depends on Gaussian-kernel design. However, designing a generalized density map for objects with diverse sizes and shapes is inherently difficult, which limits the robustness of existing methods in open and heterogeneous scenes. To address this issue, we propose SeqCount, a novel sequence-modeling framework that reformulates CAC as a sequence generation task rather than density regression. Specifically, we divide the input image into N×N patches and introduce a serialization scheme that converts point annotations into an ordered sequence of discrete patch-level count tokens. We further design an encoder–decoder architecture to model the correlation among image patches, where cross-attention guides each token prediction to the relevant visual regions and masked self-attention captures dependencies among neighboring counting tokens. By replacing density estimation with discrete sequence prediction, SeqCount avoids handcrafted kernel design and provides a more flexible and interpretable counting paradigm for objects with varying scales and shapes. Experimental results on three challenging multi-category datasets, i.e., FSCD-LVIS, UAVVIC and FSC-, and three vehicle counting datasets, i.e., TRANCOS, CARPK, and PUCPR+, demonstrate that our method performs favorably against state-of-the-art methods.

Original languageEnglish
Article number104856
JournalJournal of Visual Communication and Image Representation
Volume119
DOIs
StatePublished - Aug 2026
Externally publishedYes

Keywords

  • Class-agnostic counting
  • Object counting
  • Sequence modeling

Fingerprint

Dive into the research topics of 'SeqCount: A sequence modeling framework for class-agnostic counting'. Together they form a unique fingerprint.

Cite this