Skip to main navigation Skip to search Skip to main content

LEARNING SPATIAL-SEMANTIC FEATURES FOR ROBUST VIDEO OBJECT SEGMENTATION

  • Xin Li
  • , Deshui Miao
  • , Zhenyu He*
  • , Yaowei Wang*
  • , Huchuan Lu
  • , Ming Hsuan Yang
  • *Corresponding author for this work
  • Pengcheng Laboratory
  • Harbin Institute of Technology Shenzhen
  • Pazhou Laboratory (Huangpu)
  • Dalian University of Technology
  • University of California Merced
  • Yonsei University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Tracking and segmenting multiple similar objects with distinct or complex parts in long-term videos is particularly challenging due to the ambiguity in identifying target components and the confusion caused by occlusion, background clutter, and changes in appearance or environment over time. In this paper, we propose a robust video object segmentation framework that learns spatial-semantic features and discriminative object queries to address the above issues. Specifically, we construct a spatial-semantic block comprising a semantic embedding component and a spatial dependency modeling part for associating global semantic features and local spatial features, providing a comprehensive target representation. In addition, we develop a masked cross-attention module to generate object queries that focus on the most discriminative parts of target objects during query propagation, alleviating noise accumulation to ensure effective long-term query propagation. Extensive experimental results show that the proposed method achieves state-of-the-art performance on benchmark data sets, including the DAVIS2017 test (87.8%), YoutubeVOS 2019 (88.1%), MOSE val (74.0%), and LVOS test (73.0%), and demonstrate the effectiveness and generalization capacity of our model. The source code and trained models are released at https://github.com/yahooo-m/S3.

Original languageEnglish
Title of host publication13th International Conference on Learning Representations, ICLR 2025
PublisherInternational Conference on Learning Representations, ICLR
Pages32543-32562
Number of pages20
ISBN (Electronic)9798331320850
StatePublished - 2025
Externally publishedYes
Event13th International Conference on Learning Representations, ICLR 2025 - Singapore, Singapore
Duration: 24 Apr 202528 Apr 2025

Publication series

Name13th International Conference on Learning Representations, ICLR 2025

Conference

Conference13th International Conference on Learning Representations, ICLR 2025
Country/TerritorySingapore
CitySingapore
Period24/04/2528/04/25

Fingerprint

Dive into the research topics of 'LEARNING SPATIAL-SEMANTIC FEATURES FOR ROBUST VIDEO OBJECT SEGMENTATION'. Together they form a unique fingerprint.

Cite this