TY - GEN
T1 - A system based on sequence learning for event detection in surveillance video
AU - Fang, Xiaoyu
AU - Xia, Ziwei
AU - Su, Chi
AU - Xu, Teng
AU - Tian, Yonghong
AU - Wang, Yaowei
AU - Huang, Tiejun
PY - 2013
Y1 - 2013
N2 - Event detection in crowded surveillance videos is a challenging yet important problem. In this paper, we present our eSur (Event detection system on SURveillance video) system, which is derived from TRECVid'12 surveillance tasks. Currently, eSur attempts to detect two categories of events: 1) pair-wise events (e.g., PeopleMeet, PeopleSplitUp and Embrace); 2) action-like events (e.g., ObjectPut, CellToEar, PersonRuns and Pointing). In eSur system, we first employ people detection and tracking algorithms to locate target persons in 3D space-time domain. Then the video sequences in which target persons occur are partitioned into several spatio-temporal cubes. Visual features (i.e. cubic feature and MoSIFT) are computed over these cubes. After that, a sequence learning method, (namely SVM with dynamic time alignment kernel), is employed to infer the existence of an event for the video sequence. According to the TRECVid SED formal evaluation, eSur has yielded fairly encouraging results on TRECVid'12 dataset.
AB - Event detection in crowded surveillance videos is a challenging yet important problem. In this paper, we present our eSur (Event detection system on SURveillance video) system, which is derived from TRECVid'12 surveillance tasks. Currently, eSur attempts to detect two categories of events: 1) pair-wise events (e.g., PeopleMeet, PeopleSplitUp and Embrace); 2) action-like events (e.g., ObjectPut, CellToEar, PersonRuns and Pointing). In eSur system, we first employ people detection and tracking algorithms to locate target persons in 3D space-time domain. Then the video sequences in which target persons occur are partitioned into several spatio-temporal cubes. Visual features (i.e. cubic feature and MoSIFT) are computed over these cubes. After that, a sequence learning method, (namely SVM with dynamic time alignment kernel), is employed to infer the existence of an event for the video sequence. According to the TRECVid SED formal evaluation, eSur has yielded fairly encouraging results on TRECVid'12 dataset.
KW - Event detection
KW - sequence learning
KW - surveillance
UR - https://www.scopus.com/pages/publications/84897753118
U2 - 10.1109/ICIP.2013.6738740
DO - 10.1109/ICIP.2013.6738740
M3 - 会议稿件
AN - SCOPUS:84897753118
SN - 9781479923410
T3 - 2013 IEEE International Conference on Image Processing, ICIP 2013 - Proceedings
SP - 3587
EP - 3591
BT - 2013 IEEE International Conference on Image Processing, ICIP 2013 - Proceedings
PB - IEEE Computer Society
T2 - 2013 20th IEEE International Conference on Image Processing, ICIP 2013
Y2 - 15 September 2013 through 18 September 2013
ER -