Skip to main navigation Skip to search Skip to main content

VideoAligner: Text-driven feature decomposition for precise video–text alignment

  • Zhanzhou Feng
  • , Shunan Mao
  • , Yaowei Wang
  • , Shiliang Zhang*
  • *Corresponding author for this work
  • Peking University
  • Wuhan University
  • Harbin Institute of Technology Shenzhen
  • Peng Cheng Laboratory

Research output: Contribution to journalArticlepeer-review

Abstract

In text based video retrieval, text queries commonly convey certain objects and events the user desires to retrieve, which are more concise cues than video contents. Distinct information density between video contents and query cues leads to the difficulty of query-video alignment. To pursue a more compact video representation and accurate textual–visual feature matching, this paper introduces a novel VideoAligner to disentangle video features. VideoAligner first generates ‘object’ and ‘event’ tokens from query texts. It subsequently spots and merges visual tokens related to concepts in the query. In other words, we use ‘object’ and ‘event’ tokens to represent cues of query, which therefore supervise the disentanglement and extraction of meaningful visual features from videos. VideoAligner finally leads to compact visual tokens explicitly depicting query objects and events. Extensive experiments on three widely-used datasets demonstrate the promising performance and domain generalization capability of our method. For instance, our method shows better efficiency and consistently outperforms many recent works like ProST on three datasets. We hope to inspire future work for collaborative cross-modal learning with certain modality as guidance.

Original languageEnglish
Article number113971
JournalPattern Recognition
Volume180
DOIs
StatePublished - Dec 2026
Externally publishedYes

Keywords

  • Representation learning
  • Video signal processing
  • Visual information retrieval
  • Visualization

Fingerprint

Dive into the research topics of 'VideoAligner: Text-driven feature decomposition for precise video–text alignment'. Together they form a unique fingerprint.

Cite this