TY - GEN
T1 - Distillation-Based Multi-exit Fully Convolutional Network for Visual Tracking
AU - Ma, Ding
AU - Wu, Xiangqian
N1 - Publisher Copyright:
© 2021, Springer Nature Switzerland AG.
PY - 2021
Y1 - 2021
N2 - Obtaining a trade-off between accuracy and efficiency for a convolutional neural network is highly desired in the deep classification-based trackers. However, it is observed that existing methods make the predictions with the latest exits strategy for all the samples, making such strategy a time-consuming solution. Motivated by this, we propose a multi-exit architecture based on the principle of knowledge distillation to improve the speed of prediction by encouraging early exits to imitate later and more accurate exits. Specifically, we propose a distillation-based multi-exit fully convolutional network (FCN), named DMENet, for visual tracking. In DMENet, different types of attention mechanisms are embedded into different representation levels of FCN to capture more discriminative information. Then, three exits augment at different levels of FCN to handle the processing of a frame to stop early. The DMENet is trained offline with knowledge distillation to improve the accuracy of early exits. The confidence score of an exit is utilized to decide whether to locate the target with high confidence on this exit or continue processing the next exit. The extensive evaluation performed on OTB-100, UAV123, LaSOT and VOT2018 benchmarks demonstrate the proposed tracker outperforms state-of-the-art approaches with a high speed (36 FPS).
AB - Obtaining a trade-off between accuracy and efficiency for a convolutional neural network is highly desired in the deep classification-based trackers. However, it is observed that existing methods make the predictions with the latest exits strategy for all the samples, making such strategy a time-consuming solution. Motivated by this, we propose a multi-exit architecture based on the principle of knowledge distillation to improve the speed of prediction by encouraging early exits to imitate later and more accurate exits. Specifically, we propose a distillation-based multi-exit fully convolutional network (FCN), named DMENet, for visual tracking. In DMENet, different types of attention mechanisms are embedded into different representation levels of FCN to capture more discriminative information. Then, three exits augment at different levels of FCN to handle the processing of a frame to stop early. The DMENet is trained offline with knowledge distillation to improve the accuracy of early exits. The confidence score of an exit is utilized to decide whether to locate the target with high confidence on this exit or continue processing the next exit. The extensive evaluation performed on OTB-100, UAV123, LaSOT and VOT2018 benchmarks demonstrate the proposed tracker outperforms state-of-the-art approaches with a high speed (36 FPS).
KW - Knowledge distillation
KW - Multi-exit fully convolutional network
KW - Visual tracking
UR - https://www.scopus.com/pages/publications/85118192009
U2 - 10.1007/978-3-030-88004-0_27
DO - 10.1007/978-3-030-88004-0_27
M3 - 会议稿件
AN - SCOPUS:85118192009
SN - 9783030880033
T3 - Lecture Notes in Computer Science
SP - 329
EP - 341
BT - Pattern Recognition and Computer Vision - 4th Chinese Conference, PRCV 2021, Proceedings
A2 - Ma, Huimin
A2 - Wang, Liang
A2 - Zhang, Changshui
A2 - Wu, Fei
A2 - Tan, Tieniu
A2 - Wang, Yaonan
A2 - Lai, Jianhuang
A2 - Zhao, Yao
PB - Springer Science and Business Media Deutschland GmbH
T2 - 4th Chinese Conference on Pattern Recognition and Computer Vision, PRCV 2021
Y2 - 29 October 2021 through 1 November 2021
ER -