Abstract
In text based video retrieval, text queries commonly convey certain objects and events the user desires to retrieve, which are more concise cues than video contents. Distinct information density between video contents and query cues leads to the difficulty of query-video alignment. To pursue a more compact video representation and accurate textual–visual feature matching, this paper introduces a novel VideoAligner to disentangle video features. VideoAligner first generates ‘object’ and ‘event’ tokens from query texts. It subsequently spots and merges visual tokens related to concepts in the query. In other words, we use ‘object’ and ‘event’ tokens to represent cues of query, which therefore supervise the disentanglement and extraction of meaningful visual features from videos. VideoAligner finally leads to compact visual tokens explicitly depicting query objects and events. Extensive experiments on three widely-used datasets demonstrate the promising performance and domain generalization capability of our method. For instance, our method shows better efficiency and consistently outperforms many recent works like ProST on three datasets. We hope to inspire future work for collaborative cross-modal learning with certain modality as guidance.
| Original language | English |
|---|---|
| Article number | 113971 |
| Journal | Pattern Recognition |
| Volume | 180 |
| DOIs | |
| State | Published - Dec 2026 |
| Externally published | Yes |
Keywords
- Representation learning
- Video signal processing
- Visual information retrieval
- Visualization
Fingerprint
Dive into the research topics of 'VideoAligner: Text-driven feature decomposition for precise video–text alignment'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver