Skip to main navigation Skip to search Skip to main content

Recognizing actions using salient features

  • Liang Wang*
  • , Debin Zhao
  • *Corresponding author for this work
  • School of Computer Science and Technology, Harbin Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Towards a compact video feature representation, we propose a novel feature selection methodology for action recognition based on the saliency maps of videos. Since saliency maps measure the perceptual importance of the pixels and regions in videos, selecting features using saliency maps enables us to find a feature representation that covers the informative parts of a video. Because saliency detection is a bottom-up procedure, some appearance changes or motions that are irrelevant to actions may also be detected as salient regions. To further improve the purity of the feature representation, we prune these irrelevant salient regions using the saliency values distribution and the spatial-temporal distribution of the salient regions. Extensive experiments are conducted to demonstrate that the proposed feature selection method largely improves the performance of bag-of-video-words model on action recognition based on three different attention models including a static attention model, a motion attention model and their combination.

Original languageEnglish
Title of host publicationMMSP 2011 - IEEE International Workshop on Multimedia Signal Processing
DOIs
StatePublished - 2011
Externally publishedYes
Event3rd IEEE International Workshop on Multimedia Signal Processing, MMSP 2011 - Hangzhou, China
Duration: 17 Nov 201119 Nov 2011

Publication series

NameMMSP 2011 - IEEE International Workshop on Multimedia Signal Processing

Conference

Conference3rd IEEE International Workshop on Multimedia Signal Processing, MMSP 2011
Country/TerritoryChina
CityHangzhou
Period17/11/1119/11/11

Fingerprint

Dive into the research topics of 'Recognizing actions using salient features'. Together they form a unique fingerprint.

Cite this