Skip to main navigation Skip to search Skip to main content

Self-Chained Dynamic Context Perception to Tracking by Natural Language Specification

  • Faculty of Computing, Harbin Institute of Technology
  • School of Astronautics, Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Vision-language cross-modal learning has significantly improved Tracking by Natural Language specification (TNL). Most existing TNL methods follow a Siamese-like matching paradigm, where visual search-region features and language-query features are aligned with the aid of pre-trained image-text representations. Although such representations provide strong static semantic cues, they are often less effective in explicitly modeling target-state changes described by action-related phrases in natural language queries. As a result, dynamic linguistic cues, such as verbs and motion-related descriptions, may be insufficiently emphasized during cross-modal matching. To address this issue, we propose Self-Chained Dynamic Context Perception (SeDCP), a self-chained framework for explicit dynamic query modulation and language-guided visual refinement in TNL. Specifically, SeDCP consists of two coupled chains. First, the Forward Chain performs visual-evidence-guided dynamic query modulation by injecting trajectory-aware spatiotemporal cues into the language representation, thereby enhancing phrases that describe target-state changes. Second, the Backward Chain uses the dynamically enhanced query representation to refine visual spatiotemporal features, strengthening the alignment between language cues and target-state evolution. In addition, we introduce sequence-level matching rather than isolated pairwise matching to better exploit temporal dynamics, and design a Global-Local enhanced video Transformer to capture both long-range contextual dependencies and fine-grained target details. Extensive experiments on seven standard TNL benchmarks and an additional unseen LaSOText benchmark demonstrate that SeDCP consistently outperforms state-of-the-art methods and generalizes well to unseen categories and video characteristics.

Original languageEnglish
Pages (from-to)8819-8834
Number of pages16
JournalIEEE Transactions on Image Processing
Volume35
DOIs
StatePublished - 2026
Externally publishedYes

Keywords

  • Tracking by natural language specification
  • dynamic context perception
  • self-chained architecture

Fingerprint

Dive into the research topics of 'Self-Chained Dynamic Context Perception to Tracking by Natural Language Specification'. Together they form a unique fingerprint.

Cite this