TY - GEN
T1 - Temporality-guided Masked Image Consistency for Domain Adaptive Video Segmentation
AU - Zhang, Zunhao
AU - Chen, Zhenhao
AU - Shen, Yifan
AU - Guan, Dayan
AU - Kot, Alex
N1 - Publisher Copyright:
© 2023 IEEE.
PY - 2023
Y1 - 2023
N2 - Video semantic segmentation has witnessed substantial advancements, largely due to the vast volume of labeled training samples. Nevertheless, domain adaptive video segmentation that adapts from a labeled source domain to an unlabeled target domain remains insufficiently delved into. In this paper, we propose Temporality-guided Masked Image Consistency (TgMIC), a simple yet effective approach that leverages the concept of Masked Image Modeling (MIM) to learn semantic features in the target domain. Unlike random masking strategy applied in traditional MIM, TgMIC introduces a novel temporality-guided masking strategy that samples the mask according to the distribution of optical flow, which facilitate the learning of spatial context relations in video sequence. Specifically, TgMIC masks the patches in vision transformers where the variance of optical flow is large, as these patches are known to contain noisy estimates of optical flow. In order to learn semantic information for video segmentation, TgMIC reconstructs the predictions of original frames from the masked frames. Comprehensive tests and detailed analysis on various public datasets show that our mechanism stands out, outpacing contemporaneous techniques, while ensuring no additional time costs.
AB - Video semantic segmentation has witnessed substantial advancements, largely due to the vast volume of labeled training samples. Nevertheless, domain adaptive video segmentation that adapts from a labeled source domain to an unlabeled target domain remains insufficiently delved into. In this paper, we propose Temporality-guided Masked Image Consistency (TgMIC), a simple yet effective approach that leverages the concept of Masked Image Modeling (MIM) to learn semantic features in the target domain. Unlike random masking strategy applied in traditional MIM, TgMIC introduces a novel temporality-guided masking strategy that samples the mask according to the distribution of optical flow, which facilitate the learning of spatial context relations in video sequence. Specifically, TgMIC masks the patches in vision transformers where the variance of optical flow is large, as these patches are known to contain noisy estimates of optical flow. In order to learn semantic information for video segmentation, TgMIC reconstructs the predictions of original frames from the masked frames. Comprehensive tests and detailed analysis on various public datasets show that our mechanism stands out, outpacing contemporaneous techniques, while ensuring no additional time costs.
KW - Big Video Data
KW - Deep Learning
KW - Domain Adaptation
KW - Masked Image Modeling
KW - Semantic Segmentation
UR - https://www.scopus.com/pages/publications/85185839868
U2 - 10.1109/BigMM59094.2023.00015
DO - 10.1109/BigMM59094.2023.00015
M3 - 会议稿件
AN - SCOPUS:85185839868
T3 - Proceedings - 2023 IEEE 9th Multimedia Big Data, BigMM 2023
SP - 56
EP - 63
BT - Proceedings - 2023 IEEE 9th Multimedia Big Data, BigMM 2023
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 9th IEEE International Conference on Multimedia Big Data, BigMM 2023
Y2 - 11 December 2023 through 13 December 2023
ER -