Abstract
Despite Siamese trackers attracting much attention due to their scalability and efficiency in recent years, researchers have ignored the background appearance, which leads to their inapplicability in recognizing arbitrary target objects with various variations, especially in complex scenarios with background clutter and distractors. In this paper, we present a simple yet effective Siamese tracker, where the shifted windows multi-head self-attention is produced to learn the characteristics of a specific given target object for visual tracking. To validate the effectiveness of our proposed tracker, we use the Swin Transformer as the backbone network and introduced an auxiliary feature enhancement network. Extensive experimental results on two evaluation datasets demonstrate that the proposed tracker outperforms other baselines.
| Original language | English |
|---|---|
| Pages (from-to) | 161-164 |
| Number of pages | 4 |
| Journal | IEICE Transactions on Information and Systems |
| Volume | E107.D |
| Issue number | 1 |
| DOIs | |
| State | Published - Jan 2024 |
| Externally published | Yes |
Keywords
- Siamese network
- self-attention
- vision transformer
- visual tracking
Fingerprint
Dive into the research topics of 'Robust Visual Tracking Using Hierarchical Vision Transformer with Shifted Windows Multi-Head Self-Attention'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver