TY - GEN
T1 - Improving Stability of Gaze Target Detection in Videos
AU - Yang, Zhihao
AU - Wang, Xinming
AU - Wang, Zhiyong
AU - Xu, Qiong
AU - Xu, Xiu
AU - Liu, Honghai
N1 - Publisher Copyright:
© 2023 IEEE.
PY - 2023
Y1 - 2023
N2 - Obtaining accurate and stable results in gaze target detection is vital for the subsequent analysis of gaze meaning. However, existing image-based methods, which focus solely on enhancing accuracy, demonstrate poor stability when directly applied on videos. Especially when the video frame rate is low, even though the actual gaze target positions do not differ significantly between adjacent frames, the detected positions vary considerably. This inconsistency, stemming from the lack of temporal information, makes dynamic detection challenging and can lead to jarring outcomes. To reduce the jitter in gaze target detection in videos, we introduce an approach that integrates spatial and temporal modules to combine spatial with temporal information. Additionally, we propose a Jitter loss function to capture significant jitter and impose a strong penalty during training, which empowers our model with increased stability for dynamic detection. Based on a self-collected dataset, experiments demonstrate that our approach exhibits superior stability without compromising accuracy.
AB - Obtaining accurate and stable results in gaze target detection is vital for the subsequent analysis of gaze meaning. However, existing image-based methods, which focus solely on enhancing accuracy, demonstrate poor stability when directly applied on videos. Especially when the video frame rate is low, even though the actual gaze target positions do not differ significantly between adjacent frames, the detected positions vary considerably. This inconsistency, stemming from the lack of temporal information, makes dynamic detection challenging and can lead to jarring outcomes. To reduce the jitter in gaze target detection in videos, we introduce an approach that integrates spatial and temporal modules to combine spatial with temporal information. Additionally, we propose a Jitter loss function to capture significant jitter and impose a strong penalty during training, which empowers our model with increased stability for dynamic detection. Based on a self-collected dataset, experiments demonstrate that our approach exhibits superior stability without compromising accuracy.
KW - Dynamic Videos
KW - Gaze Target Detection
KW - Jitter
KW - Stability
KW - Temporal Information
UR - https://www.scopus.com/pages/publications/85179523248
U2 - 10.1109/IECON51785.2023.10312057
DO - 10.1109/IECON51785.2023.10312057
M3 - 会议稿件
AN - SCOPUS:85179523248
T3 - IECON Proceedings (Industrial Electronics Conference)
BT - IECON 2023 - 49th Annual Conference of the IEEE Industrial Electronics Society
PB - IEEE Computer Society
T2 - 49th Annual Conference of the IEEE Industrial Electronics Society, IECON 2023
Y2 - 16 October 2023 through 19 October 2023
ER -