Skip to main navigation Skip to search Skip to main content

Sound Bridge: Associating Egocentric and Exocentric Videos via Audio Cues

  • Sihong Huang
  • , Jiaxin Wu
  • , Xiaoyong Wei*
  • , Yi Cai*
  • , Dongmei Jiang
  • , Yaowei Wang
  • *Corresponding author for this work
  • South China University of Technology
  • Peng Cheng Laboratory
  • Hong Kong Polytechnic University
  • Sichuan University
  • Harbin Institute of Technology Shenzhen

Research output: Contribution to journalConference articlepeer-review

Abstract

Understanding human behavior and environmental information in egocentric videos is very challenging due to the invisibility of some actions (e.g., laughing and sneezing) and the local nature of the first-person view. Leveraging the corresponding exocentric video to provide global context has shown promising results. However, existing visual-to-visual and visual-to-textual Ego-Exo video alignment methods struggle with the issue that some activities may have non-visual overlap. To address this, we propose using sound as a bridge, as audio is often consistent across Ego-Exo videos. However, direct audio-to-audio alignment lacks context. Thus, we introduce two context-aware sound modules: one aligns audio with vision via a visual-audio cross-attention module, and another aligns text with sound closed caption generated by LLM. Experimental results on two Ego-Exo video association benchmarks show that each of the proposed modules enhances the state-of-the-art methods. Moreover, the proposed sound-aware egocentric or exocentric representation boosts the performance of downstream tasks, such as action recognition of exocentric videos and scene recognition of egocentric videos. The code and models can be accessed at https://github.com/shhuangcoder/SoundBridge.

Original languageEnglish
Pages (from-to)28942-28951
Number of pages10
JournalProceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition
DOIs
StatePublished - 2025
Externally publishedYes
Event2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2025 - Nashville, United States
Duration: 11 Jun 202515 Jun 2025

Keywords

  • audio attention
  • closed caption
  • cross-view video association
  • multimodal alignment

Fingerprint

Dive into the research topics of 'Sound Bridge: Associating Egocentric and Exocentric Videos via Audio Cues'. Together they form a unique fingerprint.

Cite this