Skip to main navigation Skip to search Skip to main content

Video Language Model Pretraining with Spatio-temporal Masking

  • Yue Wu
  • , Zhaobo Qi
  • , Junshu Sun
  • , Yaowei Wang
  • , Qingming Huang
  • , Shuhui Wang*
  • *Corresponding author for this work
  • Chinese Academy of Sciences
  • Pengcheng Laboratory
  • University of Chinese Academy of Sciences
  • Harbin Institute of Technology Weihai
  • Harbin Institute of Technology Shenzhen

Research output: Contribution to journalConference articlepeer-review

Abstract

The development of self-supervised video-language models based on mask learning has significantly advanced downstream video tasks. These models leverage masked reconstruction to facilitate joint learning of visual and linguistic information. However, recent study reveals that reconstructing image features yields superior downstream performance compared to video feature reconstruction. We hypothesize that this performance gap stems from the way how masking strategies influence the model's attention to temporal dynamics. To validate this hypothesis, we performed two sets of experiments that demonstrate that alignment between the masked target and the reconstruction target is crucial for self-supervised video-language learning. Based on these findings, we propose a spatio-temporal masking strategy (STM) for video-language model pretraining that operates across adjacent frames, and a decoder leverages semantic information to enhance the spatio-temporal representations of masked tokens. Thanks to the combination of masking strategy and reconstruction decoder, STM enforces the model to learn spatio-temporal feature representation comprehensively. Experiments in three video understanding downstream tasks validate the superiority of our method. Codes are available here.

Original languageEnglish
Pages (from-to)8557-8567
Number of pages11
JournalProceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition
DOIs
StatePublished - 2025
Externally publishedYes
Event2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2025 - Nashville, United States
Duration: 11 Jun 202515 Jun 2025

Keywords

  • spatio-temporal masking strategy
  • video-language models

Fingerprint

Dive into the research topics of 'Video Language Model Pretraining with Spatio-temporal Masking'. Together they form a unique fingerprint.

Cite this