Abstract
Video Corpus Moment Retrieval (VCMR) aims to retrieve the target temporal moment from a large-scale corpus of untrimmed videos given a natural-language query. A central challenge of this task lies in effectively modeling heterogeneous video modalities, such as visual content and subtitles, together with the query. Existing methods typically fuse multiple video modalities before aligning the fused representation with the query. Although effective under the full-modality setting, such tightly coupled modeling may obscure modality-specific discriminative cues and reduce robustness when modalities are noisy, imbalanced, or unavailable. To address this issue, we propose DeMUL, a new VCMR framework with Decoupled Multimodal Modeling and Unified Localization. Instead of directly fusing different video modalities, DeMUL first aligns the query with each modality separately and then integrates the resulting query-aligned modality-specific representations for temporal localization. This decoupled design preserves modality-specific information while still exploiting complementary multimodal evidence. Building on this formulation, we further introduce a unified localization module that jointly optimizes unimodal and multimodal branches in a shared prediction space, enabling consistent localization behavior across branches without requiring auxiliary distillation or additional supervision. Extensive experiments on TVR and DiDeMo demonstrate that DeMUL achieves state-of-the-art performance under the standard full-modality setting, improving the performance from 18.95 to 21.37 on TVR (+2.42, 12.8%) and from 4.77 to 5.52 on DiDeMo (+0.75, 15.7%) compared with the previous state-of-the-art method QCLPL. In addition, DeMUL maintains strong performance under modality-missing and degraded-input conditions, demonstrating superior robustness in realistic multimodal retrieval scenarios.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Circuits and Systems for Video Technology |
| DOIs | |
| State | Accepted/In press - 2026 |
| Externally published | Yes |
Keywords
- Multimodal representation learning
- Video corpus moment retrieval
- robustness to modality missingness
Fingerprint
Dive into the research topics of 'Video Corpus Moment Retrieval via Decoupled Multimodal Modeling and Unified Localization'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver