Skip to main navigation Skip to search Skip to main content

Robust Visual Tracking Using Hierarchical Vision Transformer with Shifted Windows Multi-Head Self-Attention

  • Peng Gao*
  • , Xin Yue Zhang
  • , Xiao Li Yang
  • , Jian Cheng Ni
  • , Fei Wang
  • *Corresponding author for this work
  • Qufu Normal University
  • School of Electronics and Information Engineering, Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Despite Siamese trackers attracting much attention due to their scalability and efficiency in recent years, researchers have ignored the background appearance, which leads to their inapplicability in recognizing arbitrary target objects with various variations, especially in complex scenarios with background clutter and distractors. In this paper, we present a simple yet effective Siamese tracker, where the shifted windows multi-head self-attention is produced to learn the characteristics of a specific given target object for visual tracking. To validate the effectiveness of our proposed tracker, we use the Swin Transformer as the backbone network and introduced an auxiliary feature enhancement network. Extensive experimental results on two evaluation datasets demonstrate that the proposed tracker outperforms other baselines.

Original languageEnglish
Pages (from-to)161-164
Number of pages4
JournalIEICE Transactions on Information and Systems
VolumeE107.D
Issue number1
DOIs
StatePublished - Jan 2024
Externally publishedYes

Keywords

  • Siamese network
  • self-attention
  • vision transformer
  • visual tracking

Fingerprint

Dive into the research topics of 'Robust Visual Tracking Using Hierarchical Vision Transformer with Shifted Windows Multi-Head Self-Attention'. Together they form a unique fingerprint.

Cite this