Skip to main navigation Skip to search Skip to main content

Language Supervised Multi-Camera Multi-Object Tracking

  • Faculty of Computing, Harbin Institute of Technology
  • Peng Cheng Laboratory

Research output: Contribution to journalArticlepeer-review

Abstract

Recent multi-camera multi-object tracking (MCMOT) algorithms are primarily trained using per-detection identity annotations, which are complicated to obtain. In contrast, labeling a language description per-object is a more natural and human-friendly way. In this paper, we explore MCMOT in a language-supervised manner (LS-MCMOT) and propose a novel approach LaVST, which performs language-to-vision weakly-supervised learning based on reliable pseudo-labels generated via tracklet-level cross-modality matching. In addition, we design an ID-aware projection self-correction mechanism to correct inaccurate image-to-ground projection in a self-supervised manner. The models trained with our approach exhibit promising performance in LS-MCMOT. Surprisingly, they perform favorably against state-of-the-art identity-supervised methods, especially in cross-dataset evaluation (with an average gain by 20.0% in IDF1), underscoring the potential of language annotations in MCMOT. Codes and language annotations will be available here.

Original languageEnglish
Pages (from-to)6123-6138
Number of pages16
JournalIEEE Transactions on Image Processing
Volume35
DOIs
StatePublished - 2026
Externally publishedYes

Keywords

  • Multi-camera multi-object tracking
  • language supervision
  • projection self-correction
  • tracklet-level matching

Fingerprint

Dive into the research topics of 'Language Supervised Multi-Camera Multi-Object Tracking'. Together they form a unique fingerprint.

Cite this