Skip to main navigation Skip to search Skip to main content

Collaborative Debias Strategy for Temporal Sentence Grounding in Video

  • Zhaobo Qi
  • , Yibo Yuan
  • , Xiaowen Ruan
  • , Shuhui Wang
  • , Weigang Zhang*
  • , Qingming Huang
  • *Corresponding author for this work
  • School of Computer Science and Technology, Harbin Institute of Technology
  • Peng Cheng Laboratory
  • CAS - Institute of Computing Technology
  • Chinese Academy of Sciences

Research output: Contribution to journalArticlepeer-review

Abstract

Temporal sentence grounding in video has witnessed significant advancements, but suffers from substantial dataset bias, which undermines its generalization ability. Existing debias approaches primarily concentrate on well-known distribution and linguistic biases, while overlooking the relationship among different biases, limiting their debias capability. In this work, we delve into the existence of visual bias and combinatorial bias in the widely used datasets, and introduce a collaborative debias structure that can be seamlessly integrated into present methods. It encompasses four low-capacity models, a re-label module, and a main model. Each biased model deliberately leverages bias as shortcut information to accurately perform grounding, achieved by customizing the appropriate model structure and input data format to align with the bias characteristics. During the training phase, the gradient descent direction for optimizing the main model should align with the negative gradient descent direction of the biased model that is optimized by utilizing ground truth labels. Subsequently, the re-label module introduces a gradient aggregation function, consolidating the gradient descent direction from these biased models and constructing new labels to compel the main model to effectively capture multi-modality alignment features instead of relying on shortcut contents for grounding. Finally, we design two debias structures, P-Debias and C-Debias, to exploit the independence and inclusion relationships between different types of biases. Extensive experiments on multiple span-based models over Charades-CD and ActivityNet-CD demonstrate the exceptional debias capability of our strategy (https://github.com/qzhb/CDS).

Original languageEnglish
Pages (from-to)10972-10986
Number of pages15
JournalIEEE Transactions on Circuits and Systems for Video Technology
Volume34
Issue number11
DOIs
StatePublished - 2024
Externally publishedYes

Keywords

  • Temporal sentence grounding in video
  • collaborative debias
  • combinatorial bias
  • visual bias

Fingerprint

Dive into the research topics of 'Collaborative Debias Strategy for Temporal Sentence Grounding in Video'. Together they form a unique fingerprint.

Cite this