Abstract
Vision-language cross-modal learning has significantly improved Tracking by Natural Language specification (TNL). Most existing TNL methods follow a Siamese-like matching paradigm, where visual search-region features and language-query features are aligned with the aid of pre-trained image-text representations. Although such representations provide strong static semantic cues, they are often less effective in explicitly modeling target-state changes described by action-related phrases in natural language queries. As a result, dynamic linguistic cues, such as verbs and motion-related descriptions, may be insufficiently emphasized during cross-modal matching. To address this issue, we propose Self-Chained Dynamic Context Perception (SeDCP), a self-chained framework for explicit dynamic query modulation and language-guided visual refinement in TNL. Specifically, SeDCP consists of two coupled chains. First, the Forward Chain performs visual-evidence-guided dynamic query modulation by injecting trajectory-aware spatiotemporal cues into the language representation, thereby enhancing phrases that describe target-state changes. Second, the Backward Chain uses the dynamically enhanced query representation to refine visual spatiotemporal features, strengthening the alignment between language cues and target-state evolution. In addition, we introduce sequence-level matching rather than isolated pairwise matching to better exploit temporal dynamics, and design a Global-Local enhanced video Transformer to capture both long-range contextual dependencies and fine-grained target details. Extensive experiments on seven standard TNL benchmarks and an additional unseen LaSOText benchmark demonstrate that SeDCP consistently outperforms state-of-the-art methods and generalizes well to unseen categories and video characteristics.
| Original language | English |
|---|---|
| Pages (from-to) | 8819-8834 |
| Number of pages | 16 |
| Journal | IEEE Transactions on Image Processing |
| Volume | 35 |
| DOIs | |
| State | Published - 2026 |
| Externally published | Yes |
Keywords
- Tracking by natural language specification
- dynamic context perception
- self-chained architecture
Fingerprint
Dive into the research topics of 'Self-Chained Dynamic Context Perception to Tracking by Natural Language Specification'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver