Skip to main navigation Skip to search Skip to main content

Sequential Deep Trajectory Descriptor for Action Recognition with Three-Stream CNN

  • Yemin Shi
  • , Yonghong Tian*
  • , Yaowei Wang
  • , Tiejun Huang
  • *Corresponding author for this work
  • Peking University
  • Beijing Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Learning the spatialoral representation of motion information is crucial to human action recognition. Nevertheless, most of the existing features or descriptors cannot capture motion information effectively, especially for long-term motion. To address this problem, this paper proposes a long-term motion descriptor called sequential deep trajectory descriptor (sDTD). Specifically, we project dense trajectories into two-dimensional planes, and subsequently a CNN-RNN network is employed to learn an effective representation for long-term motion. Unlike the popular two-stream ConvNets, the sDTD stream is introduced into a three-stream framework so as to identify actions from a video sequence. Consequently, this three-stream framework can simultaneously capture static spatial features, short-term motion, and long-term motion in the video. Extensive experiments were conducted on three challenging datasets: KTH, HMDB51, and UCF101. Experimental results show that our method achieves state-of-the-art performance on the KTH and UCF101 datasets, and is comparable to the state-of-the-art methods on the HMDB51 dataset.

Original languageEnglish
Article number7847353
Pages (from-to)1510-1520
Number of pages11
JournalIEEE Transactions on Multimedia
Volume19
Issue number7
DOIs
StatePublished - Jul 2017
Externally publishedYes

Keywords

  • Action recognition
  • long-term motion
  • sequential deep trajectory descriptor (sDTD)
  • three-stream framework

Fingerprint

Dive into the research topics of 'Sequential Deep Trajectory Descriptor for Action Recognition with Three-Stream CNN'. Together they form a unique fingerprint.

Cite this