Skip to main navigation Skip to search Skip to main content

Directly Locating Actions in Video with Single Frame Annotation

  • Haoran Tong*
  • , Xinyan Liu
  • , Guorong Li
  • , Laiyun Qing
  • *Corresponding author for this work
  • University of Chinese Academy of Sciences

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

We propose a novel method for point-supervised action localization.Differs from the common practice of locating actions by first categorizing each video frame, our method directly predicts actions’ positions and length. Specifically, point-supervised action localization is achieved by a series of fully supervised action location iteratively. In each iteration, the input video are used as input tokens and fed into a transformer, where the encoder extracts global context of the clips, and the decoder generates queries containing information for action localization. Three MLP heads are built on each query to obtain the probability, the center, and the length of each action instance respectively. Experiments on three popular datasets prove the potential of our method.

Original languageEnglish
Title of host publicationICMR 2024-Proceedings of the 14th Annual ACM International Conference on Multimedia Retrieval
PublisherAssociation for Computing Machinery, Inc
Pages1135-1139
Number of pages5
ISBN (Electronic)9798400706028
DOIs
StatePublished - 7 Jun 2024
Externally publishedYes
Event14th Annual ACM International Conference on Multimedia Retrieval, ICMR 2024 - Phuket, Thailand
Duration: 10 Jun 202414 Jun 2024

Publication series

NameICMR 2024 - Proceedings of the 2024 International Conference on Multimedia Retrieval

Conference

Conference14th Annual ACM International Conference on Multimedia Retrieval, ICMR 2024
Country/TerritoryThailand
CityPhuket
Period10/06/2414/06/24

Keywords

  • Point-level supervision
  • Temporal action detection

Fingerprint

Dive into the research topics of 'Directly Locating Actions in Video with Single Frame Annotation'. Together they form a unique fingerprint.

Cite this