Abstract
Video salient object detection (VSOD) aims to identify and highlight the most visually compelling and motion-related elements within video sequences, which serves as a crucial preprocessing step for intelligent video analysis. However, due to inadequate spatiotemporal cross-modal feature fusion and suboptimal capture of salient structure information, existing detection methods exhibit poor performance in numerous complex scenes. Such scenes typically involve moving objects but not salient in the background or rapid motion changes in foreground objects. To address this issue, we propose a multi-graph mutual learning network with cross-modal feature fusion for VSOD. Specifically, we design a cross-attention module (CAM) to effectively fuse spatiotemporal modal features. And we devise a multi-scale feature fusion module (MFFM) to fully integrate multi-scale features from different feature extraction layers. Finally, we propose a multi-graph mutual learning network (MGMLN) to improve the integrity and continuity of object structural information. Extensive experiments were conducted on four commonly used test datasets for VSOD. Our proposed method demonstrates the capability to accurately predict the most salient objects and maintain coherent details within complex dynamic visual scenes, when benchmarked against 21 state-of-the-art (SOTA) VSOD models.
| Original language | English |
|---|---|
| Article number | 105485 |
| Journal | Digital Signal Processing: A Review Journal |
| Volume | 168 |
| DOIs | |
| State | Published - Jan 2026 |
| Externally published | Yes |
Keywords
- Cross-modal feature fusion
- Graph neural network
- Multi-graph mutual learning
- Video salient object detection
Fingerprint
Dive into the research topics of 'Multi-graph mutual learning network with cross-modal feature fusion for video salient object detection'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver