Skip to main navigation Skip to search Skip to main content

Temporality-guided Masked Image Consistency for Domain Adaptive Video Segmentation

  • Zunhao Zhang
  • , Zhenhao Chen
  • , Yifan Shen
  • , Dayan Guan
  • , Alex Kot
  • Hong Kong Industrial Artificial Intelligence and Robotics Center
  • Mohamed Bin Zayed University of Artificial Intelligence
  • Nanyang Technological University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Video semantic segmentation has witnessed substantial advancements, largely due to the vast volume of labeled training samples. Nevertheless, domain adaptive video segmentation that adapts from a labeled source domain to an unlabeled target domain remains insufficiently delved into. In this paper, we propose Temporality-guided Masked Image Consistency (TgMIC), a simple yet effective approach that leverages the concept of Masked Image Modeling (MIM) to learn semantic features in the target domain. Unlike random masking strategy applied in traditional MIM, TgMIC introduces a novel temporality-guided masking strategy that samples the mask according to the distribution of optical flow, which facilitate the learning of spatial context relations in video sequence. Specifically, TgMIC masks the patches in vision transformers where the variance of optical flow is large, as these patches are known to contain noisy estimates of optical flow. In order to learn semantic information for video segmentation, TgMIC reconstructs the predictions of original frames from the masked frames. Comprehensive tests and detailed analysis on various public datasets show that our mechanism stands out, outpacing contemporaneous techniques, while ensuring no additional time costs.

Original languageEnglish
Title of host publicationProceedings - 2023 IEEE 9th Multimedia Big Data, BigMM 2023
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages56-63
Number of pages8
ISBN (Electronic)9798350360004
DOIs
StatePublished - 2023
Externally publishedYes
Event9th IEEE International Conference on Multimedia Big Data, BigMM 2023 - Hybrid, Laguna Hills, United States
Duration: 11 Dec 202313 Dec 2023

Publication series

NameProceedings - 2023 IEEE 9th Multimedia Big Data, BigMM 2023

Conference

Conference9th IEEE International Conference on Multimedia Big Data, BigMM 2023
Country/TerritoryUnited States
CityHybrid, Laguna Hills
Period11/12/2313/12/23

Keywords

  • Big Video Data
  • Deep Learning
  • Domain Adaptation
  • Masked Image Modeling
  • Semantic Segmentation

Fingerprint

Dive into the research topics of 'Temporality-guided Masked Image Consistency for Domain Adaptive Video Segmentation'. Together they form a unique fingerprint.

Cite this