Abstract
Predicting long-term scanpaths in panoramic videos requires effective fusion of multimodal inputs, including visual content and historical gaze sequences. Existing methods typically process these modalities independently, ignoring their intrinsic temporal-spatial correlations and thereby limiting the fidelity of multimodal distribution modeling. This paper introduces a unified framework that explicitly aligns visual and gaze modalities in both time and space. A hybrid-granularity Transformer is proposed to jointly encode global semantic structures and local fine-grained dynamics, enabling more accurate long-term dependency modeling. To further enhance cross-modal fusion, a contrastive learning strategy is employed to improve the alignment of visual features and scanpath representations. For scanpath generation, a lightweight optimization-based sampler-guided by a physics-inspired proxy viewer-is integrated to produce smooth and realistic gaze trajectories without relying on heuristic sampling. Evaluations on VRW-23 and CVPR-18 demonstrate consistent state-of-the-art performance, confirming the effectiveness of the proposed temporal-spatial multimodal alignment and hybrid-granularity architecture.
| Original language | English |
|---|---|
| Article number | 113480 |
| Journal | Pattern Recognition |
| Volume | 179 |
| DOIs | |
| State | Published - Nov 2026 |
| Externally published | Yes |
Keywords
- Contrastive learning
- Multimodal learning
- Panoramic video
- Scanpath prediction
- Transformer architecture
Fingerprint
Dive into the research topics of 'Gradient descent-driven sampling for multimodal long-term scanpath prediction in panoramic videos'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver