Abstract
Recent multi-camera multi-object tracking (MCMOT) algorithms are primarily trained using per-detection identity annotations, which are complicated to obtain. In contrast, labeling a language description per-object is a more natural and human-friendly way. In this paper, we explore MCMOT in a language-supervised manner (LS-MCMOT) and propose a novel approach LaVST, which performs language-to-vision weakly-supervised learning based on reliable pseudo-labels generated via tracklet-level cross-modality matching. In addition, we design an ID-aware projection self-correction mechanism to correct inaccurate image-to-ground projection in a self-supervised manner. The models trained with our approach exhibit promising performance in LS-MCMOT. Surprisingly, they perform favorably against state-of-the-art identity-supervised methods, especially in cross-dataset evaluation (with an average gain by 20.0% in IDF1), underscoring the potential of language annotations in MCMOT. Codes and language annotations will be available here.
| Original language | English |
|---|---|
| Pages (from-to) | 6123-6138 |
| Number of pages | 16 |
| Journal | IEEE Transactions on Image Processing |
| Volume | 35 |
| DOIs | |
| State | Published - 2026 |
| Externally published | Yes |
Keywords
- Multi-camera multi-object tracking
- language supervision
- projection self-correction
- tracklet-level matching
Fingerprint
Dive into the research topics of 'Language Supervised Multi-Camera Multi-Object Tracking'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver