TY - GEN
T1 - STGA-Net
T2 - 2023 IEEE International Conference on Multimedia and Expo Workshops, ICMEW 2023
AU - Tian, Xiaoyan
AU - Jin, Ye
AU - Zhang, Zhao
AU - Liu, Peng
AU - Tang, Xianglong
N1 - Publisher Copyright:
© 2023 IEEE.
PY - 2023
Y1 - 2023
N2 - Temporal action segmentation aims at dense labeling of video frames with a series of action classes in long and untrimmed videos. However, previous methods heavily rely on generating an initial prediction with temporal convolutional layers and refining the predictions over the following stages based on RGB features. This results in the lack of an explicit action segment transition rule and loss of high-level semantic information, such as complex spatial-temporal correlation among the human joints between frames. Therefore, we present a spatial-temporal graph attention network (STGA-Net) for skeleton-based temporal action segmentation. In particular, we propose a spatial-temporal attentive block for prediction generation, which adapts an encoder-decoder architecture, where both the encoder and decoder contain various graph spatial-temporal attention blocks to model the dynamic and non-linear correlation among joints. Experiments on three challenging datasets (PKU-MMD, HuGaDB, and LARa) demonstrate that the performance of our STGA-Net exceeds that of the state-of-the-art and alleviates over-segmentation and ambiguous boundary errors to a large degree.
AB - Temporal action segmentation aims at dense labeling of video frames with a series of action classes in long and untrimmed videos. However, previous methods heavily rely on generating an initial prediction with temporal convolutional layers and refining the predictions over the following stages based on RGB features. This results in the lack of an explicit action segment transition rule and loss of high-level semantic information, such as complex spatial-temporal correlation among the human joints between frames. Therefore, we present a spatial-temporal graph attention network (STGA-Net) for skeleton-based temporal action segmentation. In particular, we propose a spatial-temporal attentive block for prediction generation, which adapts an encoder-decoder architecture, where both the encoder and decoder contain various graph spatial-temporal attention blocks to model the dynamic and non-linear correlation among joints. Experiments on three challenging datasets (PKU-MMD, HuGaDB, and LARa) demonstrate that the performance of our STGA-Net exceeds that of the state-of-the-art and alleviates over-segmentation and ambiguous boundary errors to a large degree.
KW - Skeleton-based temporal action segmentation
KW - ambiguous boundary
KW - attention
KW - over-segmentation
KW - spatial-temporal correlation
UR - https://www.scopus.com/pages/publications/85172301942
U2 - 10.1109/ICMEW59549.2023.00044
DO - 10.1109/ICMEW59549.2023.00044
M3 - 会议稿件
AN - SCOPUS:85172301942
T3 - Proceedings - 2023 IEEE International Conference on Multimedia and Expo Workshops, ICMEW 2023
SP - 218
EP - 223
BT - Proceedings - 2023 IEEE International Conference on Multimedia and Expo Workshops, ICMEW 2023
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 10 July 2023 through 14 July 2023
ER -