Skip to main navigation Skip to search Skip to main content

Multi-graph mutual learning network with cross-modal feature fusion for video salient object detection

  • School of Electronics and Information Engineering, Harbin Institute of Technology
  • Harbin Engineering University

Research output: Contribution to journalArticlepeer-review

Abstract

Video salient object detection (VSOD) aims to identify and highlight the most visually compelling and motion-related elements within video sequences, which serves as a crucial preprocessing step for intelligent video analysis. However, due to inadequate spatiotemporal cross-modal feature fusion and suboptimal capture of salient structure information, existing detection methods exhibit poor performance in numerous complex scenes. Such scenes typically involve moving objects but not salient in the background or rapid motion changes in foreground objects. To address this issue, we propose a multi-graph mutual learning network with cross-modal feature fusion for VSOD. Specifically, we design a cross-attention module (CAM) to effectively fuse spatiotemporal modal features. And we devise a multi-scale feature fusion module (MFFM) to fully integrate multi-scale features from different feature extraction layers. Finally, we propose a multi-graph mutual learning network (MGMLN) to improve the integrity and continuity of object structural information. Extensive experiments were conducted on four commonly used test datasets for VSOD. Our proposed method demonstrates the capability to accurately predict the most salient objects and maintain coherent details within complex dynamic visual scenes, when benchmarked against 21 state-of-the-art (SOTA) VSOD models.

Original languageEnglish
Article number105485
JournalDigital Signal Processing: A Review Journal
Volume168
DOIs
StatePublished - Jan 2026
Externally publishedYes

Keywords

  • Cross-modal feature fusion
  • Graph neural network
  • Multi-graph mutual learning
  • Video salient object detection

Fingerprint

Dive into the research topics of 'Multi-graph mutual learning network with cross-modal feature fusion for video salient object detection'. Together they form a unique fingerprint.

Cite this