TY - GEN
T1 - LEARNING SPATIAL-SEMANTIC FEATURES FOR ROBUST VIDEO OBJECT SEGMENTATION
AU - Li, Xin
AU - Miao, Deshui
AU - He, Zhenyu
AU - Wang, Yaowei
AU - Lu, Huchuan
AU - Yang, Ming Hsuan
N1 - Publisher Copyright:
© 2025 13th International Conference on Learning Representations, ICLR 2025. All rights reserved.
PY - 2025
Y1 - 2025
N2 - Tracking and segmenting multiple similar objects with distinct or complex parts in long-term videos is particularly challenging due to the ambiguity in identifying target components and the confusion caused by occlusion, background clutter, and changes in appearance or environment over time. In this paper, we propose a robust video object segmentation framework that learns spatial-semantic features and discriminative object queries to address the above issues. Specifically, we construct a spatial-semantic block comprising a semantic embedding component and a spatial dependency modeling part for associating global semantic features and local spatial features, providing a comprehensive target representation. In addition, we develop a masked cross-attention module to generate object queries that focus on the most discriminative parts of target objects during query propagation, alleviating noise accumulation to ensure effective long-term query propagation. Extensive experimental results show that the proposed method achieves state-of-the-art performance on benchmark data sets, including the DAVIS2017 test (87.8%), YoutubeVOS 2019 (88.1%), MOSE val (74.0%), and LVOS test (73.0%), and demonstrate the effectiveness and generalization capacity of our model. The source code and trained models are released at https://github.com/yahooo-m/S3.
AB - Tracking and segmenting multiple similar objects with distinct or complex parts in long-term videos is particularly challenging due to the ambiguity in identifying target components and the confusion caused by occlusion, background clutter, and changes in appearance or environment over time. In this paper, we propose a robust video object segmentation framework that learns spatial-semantic features and discriminative object queries to address the above issues. Specifically, we construct a spatial-semantic block comprising a semantic embedding component and a spatial dependency modeling part for associating global semantic features and local spatial features, providing a comprehensive target representation. In addition, we develop a masked cross-attention module to generate object queries that focus on the most discriminative parts of target objects during query propagation, alleviating noise accumulation to ensure effective long-term query propagation. Extensive experimental results show that the proposed method achieves state-of-the-art performance on benchmark data sets, including the DAVIS2017 test (87.8%), YoutubeVOS 2019 (88.1%), MOSE val (74.0%), and LVOS test (73.0%), and demonstrate the effectiveness and generalization capacity of our model. The source code and trained models are released at https://github.com/yahooo-m/S3.
UR - https://www.scopus.com/pages/publications/105010195135
M3 - 会议稿件
AN - SCOPUS:105010195135
T3 - 13th International Conference on Learning Representations, ICLR 2025
SP - 32543
EP - 32562
BT - 13th International Conference on Learning Representations, ICLR 2025
PB - International Conference on Learning Representations, ICLR
T2 - 13th International Conference on Learning Representations, ICLR 2025
Y2 - 24 April 2025 through 28 April 2025
ER -