Skip to main navigation Skip to search Skip to main content

Language-Guided Graph Representation Learning for Video Summarization

  • Harbin Institute of Technology
  • Peking University
  • Peng Cheng Laboratory

Research output: Contribution to journalArticlepeer-review

Abstract

With the rapid growth of video content on social media, video summarization has become a crucial task in multimedia processing. However, existing methods face challenges in capturing global dependencies in video content and accommodating multimodal user customization. Moreover, temporal proximity between video frames does not always correspond to semantic proximity. To tackle these challenges, we propose a novel Language-guided Graph Representation Learning Network (LGRLN) for video summarization. Specifically, we introduce a video graph generator that converts video frames into a structured graph to preserve temporal order and contextual dependencies. By constructing forward, backward and undirected graphs, the video graph generator effectively preserves the sequentiality and contextual relationships of video content. We designed an intra-graph relational reasoning module with a dual-threshold graph convolution mechanism, which distinguishes semantically relevant frames from irrelevant ones between nodes. Additionally, our proposed language-guided cross-modal embedding module generates video summaries with specific textual descriptions. We model the summary generation output as a mixture of Bernoulli distribution and solve it with the EM algorithm. Experimental results show that our method outperforms existing approaches across multiple benchmarks. Moreover, we proposed LGRLN reduces inference time and model parameters by 87.8% and 91.7%, respectively.

Original languageEnglish
Pages (from-to)3216-3232
Number of pages17
JournalIEEE Transactions on Pattern Analysis and Machine Intelligence
Volume48
Issue number3
DOIs
StatePublished - 2026

Keywords

  • Graph representation learning
  • query suggestion
  • video summarization

Fingerprint

Dive into the research topics of 'Language-Guided Graph Representation Learning for Video Summarization'. Together they form a unique fingerprint.

Cite this