Skip to main navigation Skip to search Skip to main content

Intra- and Inter-modal Multilinear Pooling with Multitask Learning for Video Grounding

  • Zhou Yu
  • , Yijun Song
  • , Jun Yu*
  • , Meng Wang
  • , Qingming Huang
  • *Corresponding author for this work
  • Hangzhou Dianzi University
  • Hefei University of Technology
  • University of Chinese Academy of Sciences

Research output: Contribution to journalArticlepeer-review

Abstract

Video grounding aims to temporally localize an action in an untrimmed video referred to by a query in natural language, which plays an important role in fine-grained video understanding. Given temporal proposals of limited granularity, the task is challenging that it requires fusing multi-modal features from questions and videos effectively, and localizing the referred action accurately. For multimodal feature fusion, we present an Intra- and Inter-modal Multilinear pooling (IIM) model to effectively combine the multi-modal features with considering both the intra- and inter-modal feature interactions. Compared to existing multimodal fusion models, IIM can capture high-order interactions and is more capable for modeling temporal information of videos. For action localization, we propose a simple yet effective multi-task learning framework to simultaneously predict the action label, alignment score and refined location in an end-to-end manner. Experimental results on real-world TaCoS and Charades-STA datasets demonstrate the superiority of the proposed approach over existing state-of-the-art methods.

Original languageEnglish
Pages (from-to)1863-1879
Number of pages17
JournalNeural Processing Letters
Volume52
Issue number3
DOIs
StatePublished - Dec 2020
Externally publishedYes

Keywords

  • Deep learning
  • Multimedia data analysis
  • Multimodal learning
  • Video grounding

Fingerprint

Dive into the research topics of 'Intra- and Inter-modal Multilinear Pooling with Multitask Learning for Video Grounding'. Together they form a unique fingerprint.

Cite this