Skip to main navigation Skip to search Skip to main content

Video Corpus Moment Retrieval via Decoupled Multimodal Modeling and Unified Localization

  • Yong Yang*
  • , Meng Liu
  • , Xuemeng Song
  • , Na Zheng
  • , Ke Lv
  • , Weili Guan
  • *Corresponding author for this work
  • School of Information Science and Technology, Harbin Institute of Technology Shenzhen
  • Shandong University
  • Southern University of Science and Technology
  • National University of Singapore
  • University of Chinese Academy of Sciences
  • Peng Cheng Laboratory

Research output: Contribution to journalArticlepeer-review

Abstract

Video Corpus Moment Retrieval (VCMR) aims to retrieve the target temporal moment from a large-scale corpus of untrimmed videos given a natural-language query. A central challenge of this task lies in effectively modeling heterogeneous video modalities, such as visual content and subtitles, together with the query. Existing methods typically fuse multiple video modalities before aligning the fused representation with the query. Although effective under the full-modality setting, such tightly coupled modeling may obscure modality-specific discriminative cues and reduce robustness when modalities are noisy, imbalanced, or unavailable. To address this issue, we propose DeMUL, a new VCMR framework with Decoupled Multimodal Modeling and Unified Localization. Instead of directly fusing different video modalities, DeMUL first aligns the query with each modality separately and then integrates the resulting query-aligned modality-specific representations for temporal localization. This decoupled design preserves modality-specific information while still exploiting complementary multimodal evidence. Building on this formulation, we further introduce a unified localization module that jointly optimizes unimodal and multimodal branches in a shared prediction space, enabling consistent localization behavior across branches without requiring auxiliary distillation or additional supervision. Extensive experiments on TVR and DiDeMo demonstrate that DeMUL achieves state-of-the-art performance under the standard full-modality setting, improving the performance from 18.95 to 21.37 on TVR (+2.42, 12.8%) and from 4.77 to 5.52 on DiDeMo (+0.75, 15.7%) compared with the previous state-of-the-art method QCLPL. In addition, DeMUL maintains strong performance under modality-missing and degraded-input conditions, demonstrating superior robustness in realistic multimodal retrieval scenarios.

Original languageEnglish
JournalIEEE Transactions on Circuits and Systems for Video Technology
DOIs
StateAccepted/In press - 2026
Externally publishedYes

Keywords

  • Multimodal representation learning
  • Video corpus moment retrieval
  • robustness to modality missingness

Fingerprint

Dive into the research topics of 'Video Corpus Moment Retrieval via Decoupled Multimodal Modeling and Unified Localization'. Together they form a unique fingerprint.

Cite this