Skip to main navigation Skip to search Skip to main content

ODVTrack: Only Decoder for Visual Tracking with Multiple Templates

  • Gu Geng
  • , Di Yuan*
  • , Xuyang Li*
  • , Qiao Liu
  • , Xiaojun Chang
  • , Zhenyu He
  • *Corresponding author for this work
  • Guangzhou Institute of Technology
  • Chongqing Normal University
  • University of Technology Sydney
  • Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Leveraging multiple templates across video frames has become a popular direction in visual object tracking to enhance robustness against target appearance variations. However, existing multi-template tracking methods suffer from a significant increase in computational complexity as the number of templates grows, which severely degrades inference speed. In this work, we propose a novel tracking framework based on the only-decoder paradigm, which redefines tracking as a template-guided feature extraction process. To this end, we propose a unified attention mechanism that integrates self-attention and cross-attention, enabling the model to efficiently extract discriminative features from the search region while incorporating guidance from multiple templates. This design ensures that the growth in complexity slows as the number of templates increases, thus dramatically improving inference efficiency. To further enhance runtime performance, we incorporate a Key–Value caching mechanism to eliminate redundant computations associated with template features during inference. In addition, we employ a Kalman filter to model target motion across video frames, refining the tracker’s predictions and further improving tracking accuracy. Extensive experiments show that the tracking framework we proposed has a significant advantage in speed compared to other multi-template tracking frameworks, and it also maintains very competitive accuracy on multiple datasets.

Original languageEnglish
JournalIEEE Transactions on Circuits and Systems for Video Technology
DOIs
StateAccepted/In press - 2026
Externally publishedYes

Keywords

  • Kalman filter
  • multiple templates
  • only-decoder
  • visual tracking

Fingerprint

Dive into the research topics of 'ODVTrack: Only Decoder for Visual Tracking with Multiple Templates'. Together they form a unique fingerprint.

Cite this