TY - GEN
T1 - Human action recognition based on 3D SIFT and LDA model
AU - Liu, Ping
AU - Wang, Jin
AU - She, Mary
AU - Liu, Honghai
PY - 2011
Y1 - 2011
N2 - How to recognize human action from videos captured by modern cameras efficiently and effectively is a challenge in real applications. Traditional methods which need professional analysts are facing a bottleneck because of their shortcomings. To cope with the disadvantage, methods based on computer vision techniques, without or with only a few human interventions, have been proposed to analyse human actions in videos automatically. This paper provides a method combining the three dimensional Scale Invariant Feature Transform (SIFT) detector and the Latent Dirichlet Allocation (LDA) model for human motion analysis. To represent videos effectively and robustly, we extract the 3D SIFT descriptor around each interest point, which is sampled densely from 3D Space-time video volumes. After obtaining the representation of each video frame, the LDA model is adopted to discover the underlying structure-the categorization of human actions in the collection of videos. Public available standard datasets are used to test our method. The concluding part discusses the research challenges and future directions.
AB - How to recognize human action from videos captured by modern cameras efficiently and effectively is a challenge in real applications. Traditional methods which need professional analysts are facing a bottleneck because of their shortcomings. To cope with the disadvantage, methods based on computer vision techniques, without or with only a few human interventions, have been proposed to analyse human actions in videos automatically. This paper provides a method combining the three dimensional Scale Invariant Feature Transform (SIFT) detector and the Latent Dirichlet Allocation (LDA) model for human motion analysis. To represent videos effectively and robustly, we extract the 3D SIFT descriptor around each interest point, which is sampled densely from 3D Space-time video volumes. After obtaining the representation of each video frame, the LDA model is adopted to discover the underlying structure-the categorization of human actions in the collection of videos. Public available standard datasets are used to test our method. The concluding part discusses the research challenges and future directions.
KW - 3D SIFT
KW - Human action recognition
KW - Latent Dirichlet Allocation
UR - https://www.scopus.com/pages/publications/79961220107
U2 - 10.1109/RIISS.2011.5945790
DO - 10.1109/RIISS.2011.5945790
M3 - 会议稿件
AN - SCOPUS:79961220107
SN - 9781424498840
T3 - IEEE SSCI 2011: Symposium Series on Computational Intelligence - RIISS 2011: 2011 IEEE Workshop on Robotic Intelligence in Informationally Structured Space
SP - 12
EP - 17
BT - IEEE SSCI 2011
T2 - Symposium Series on Computational Intelligence, IEEE SSCI 2011 - 2011 IEEE Workshop on Robotic Intelligence in Informationally Structured Space, RIISS 2011
Y2 - 11 April 2011 through 15 April 2011
ER -