Skip to main navigation Skip to search Skip to main content

Attention-driven action retrieval with DTW-based 3D descriptor matching

  • Rongrong Ji*
  • , Xiaoshuai Sun
  • , Hongxun Yao
  • , Pengfei Xu
  • , Tianqiang Liu
  • *Corresponding author for this work
  • Harbin Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

From visual perception viewpoint, actions in videos can capture high-level semantics for video content understanding and retrieval. However, action-level video retrieval meets great challenges, due to the interferences from global motions or concurrent actions, and the difficulties in robust action describing and matching. This paper presents a content-based action retrieval framework to enable effective search of near-duplicated actions in large-scale video database. Firstly, we present an attention shift model to distill and partition human-concerned saliency actions from global motions and concurrent actions. Secondly, to characterize each saliency action, we extract 3D-SIFT descriptor within its spatial-temporal region, which is robust against rotation, scale, and view point variances. Finally, action similarity is measured using Dynamic Time Warping (DTW) distance to offer tolerance for action duration variance and partial motion missing. Search efficiency in large-scale dataset is achieved by hierarchical descriptor indexing and approximate nearest-neighbor search. In validation, we present a prototype system VILAR to facilitate action search within "Friends" soap operas with excellent accuracy, efficiency, and human perception revealing ability.

Original languageEnglish
Title of host publicationMM'08 - Proceedings of the 2008 ACM International Conference on Multimedia, with co-located Symposium and Workshops
Pages619-622
Number of pages4
DOIs
StatePublished - 2008
Event16th ACM International Conference on Multimedia, MM '08 - Vancouver, BC, Canada
Duration: 26 Oct 200831 Oct 2008

Publication series

NameMM'08 - Proceedings of the 2008 ACM International Conference on Multimedia, with co-located Symposium and Workshops

Conference

Conference16th ACM International Conference on Multimedia, MM '08
Country/TerritoryCanada
CityVancouver, BC
Period26/10/0831/10/08

Keywords

  • 3D-sift
  • Action retrieval
  • Attention shift
  • Dynamic time warping
  • Video content analysis

Fingerprint

Dive into the research topics of 'Attention-driven action retrieval with DTW-based 3D descriptor matching'. Together they form a unique fingerprint.

Cite this