Abstract
Recently, significant progress has been made in text-video retrieval by adapting CLIP to the text-video domain. To mitigate the computational overhead of full fine-tuning, current studies have shifted focus toward parameter-efficient fine-tuning strategies, such as Adapters, Prompts, and LoRA. However, most methods generally obtain the global video representation by frame-level aggregation, i.e., pooling operation over [CLS] token in each frame, which fails to capture the local details. Meanwhile, these methods process all visual tokens indiscriminately when learning video representations, overlooking the interference introduced by redundant tokens. To tackle these issues, this paper proposes a Dual-Attention Video Representation Learning (DAVRL) framework which synergistically combines dual-attention based global-local interactions with token selection. Specifically, we incorporate a dual-attention global-local interaction paradigm and formulate a progressive cross-frame interaction module, which implement global interactions between well-designed learnable global video tokens and all visual tokens, and progressive local interactions from single-frame to multi-frame scales. Moreover, we propose a semantic-aware token selection module to prune the spatial-temporal redundant tokens, under the guidance of both global and local semantics. As such, our DAVRL framework can learn more informative yet discriminative global video representations while mitigating the interference of redundant tokens. Extensive experiments demonstrate the superiority of our DAVRL on diverse datasets with only 0.32% tunable parameters.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Multimedia |
| DOIs | |
| State | Accepted/In press - 2026 |
| Externally published | Yes |
Keywords
- Text-video retrieval
- parameter-efficient fine-tuning
- video representation learning
Fingerprint
Dive into the research topics of 'Dual-Attention Video Representation Learning for Parameter Efficient Text-Video Retrieval'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver