Skip to main navigation Skip to search Skip to main content

Gradient descent-driven sampling for multimodal long-term scanpath prediction in panoramic videos

  • Tianming Zhou
  • , Yulong Cheng
  • , Kanglong Fan
  • , Youneng Bao
  • , Mu Li*
  • *Corresponding author for this work
  • Harbin Institute of Technology Shenzhen
  • City University of Hong Kong
  • Shenzhen University

Research output: Contribution to journalArticlepeer-review

Abstract

Predicting long-term scanpaths in panoramic videos requires effective fusion of multimodal inputs, including visual content and historical gaze sequences. Existing methods typically process these modalities independently, ignoring their intrinsic temporal-spatial correlations and thereby limiting the fidelity of multimodal distribution modeling. This paper introduces a unified framework that explicitly aligns visual and gaze modalities in both time and space. A hybrid-granularity Transformer is proposed to jointly encode global semantic structures and local fine-grained dynamics, enabling more accurate long-term dependency modeling. To further enhance cross-modal fusion, a contrastive learning strategy is employed to improve the alignment of visual features and scanpath representations. For scanpath generation, a lightweight optimization-based sampler-guided by a physics-inspired proxy viewer-is integrated to produce smooth and realistic gaze trajectories without relying on heuristic sampling. Evaluations on VRW-23 and CVPR-18 demonstrate consistent state-of-the-art performance, confirming the effectiveness of the proposed temporal-spatial multimodal alignment and hybrid-granularity architecture.

Original languageEnglish
Article number113480
JournalPattern Recognition
Volume179
DOIs
StatePublished - Nov 2026
Externally publishedYes

Keywords

  • Contrastive learning
  • Multimodal learning
  • Panoramic video
  • Scanpath prediction
  • Transformer architecture

Fingerprint

Dive into the research topics of 'Gradient descent-driven sampling for multimodal long-term scanpath prediction in panoramic videos'. Together they form a unique fingerprint.

Cite this