Skip to main navigation Skip to search Skip to main content

Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning

  • Jun Li
  • , Jinpeng Wang*
  • , Chaolei Tan
  • , Niu Lian
  • , Long Chen
  • , Yaowei Wang
  • , Min Zhang
  • , Shu Tao Xia
  • , Bin Chen
  • *Corresponding author for this work
  • Harbin Institute of Technology Shenzhen
  • Tsinghua University
  • Hong Kong University of Science and Technology
  • Peng Cheng Laboratory

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Partially Relevant Video Retrieval (PRVR) addresses the critical challenge of matching untrimmed videos with text queries describing only partial content. Existing methods suffer from geometric distortion in Euclidean space that sometimes misrepresents the intrinsic hierarchical structure of videos and overlooks certain hierarchical semantics, ultimately leading to suboptimal temporal modeling. To address this issue, we propose the first hyperbolic modeling framework for PRVR, namely HLFormer, which leverages hyperbolic space learning to compensate for the suboptimal hierarchical modeling capabilities of Euclidean space. Specifically, HLFormer integrates the Lorentz Attention Block and Euclidean Attention Block to encode video embeddings in hybrid spaces, using the Mean-Guided Adaptive Interaction Module to dynamically fuse features. Additionally, we introduce a Partial Order Preservation Loss to enforce 'text ≺video' hierarchy through Lorentzian cone constraints. This approach further enhances cross-modal matching by reinforcing partial relevance between video content and text queries. Extensive experiments show that HLFormer outperforms state-of-the-art methods. Code is released at https://github.com/lijun2005/ICCV25-HLFormer.

Original languageEnglish
Title of host publicationProceedings - 2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages23074-23084
Number of pages11
ISBN (Electronic)9798331587758
DOIs
StatePublished - 2025
Externally publishedYes
Event2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025 - Honolulu, United States
Duration: 19 Oct 202523 Oct 2025

Publication series

NameProceedings of the IEEE International Conference on Computer Vision
ISSN (Print)1550-5499
ISSN (Electronic)2380-7504

Conference

Conference2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
Country/TerritoryUnited States
CityHonolulu
Period19/10/2523/10/25

Keywords

  • cross-modal retrieval
  • hyperbolic learning
  • partially relevant video retrieval
  • video representation learning
  • video-text retrieval

Fingerprint

Dive into the research topics of 'Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning'. Together they form a unique fingerprint.

Cite this