Abstract
Action recognition has attracted growing interest recently. It suffers from the problem that complex and diverse environments may disturb the extraction of action features. Existing methods propose to explore the temporal associations to alleviate the issue. However, they cannot handle long-range frames, and the rigid techniques are powerless against the differences caused by the deformation of the actors. To this end, we propose the Actor-Aware Alignment Network (A$^{3}$Net), which helps locate the action region. Specifically, through the intra-snippet correction, we afford the local segment alignment frames. The inter-snippet is designed to rectify the results, avoiding the occlusion situation that may appear in the local snippet. In addition, we consider intra-alignment short-range adjustive frames and long-range context frames between different snippets, which allows our A$^{3}$Net network to achieve the effect of focusing on long-range frame information. Multiple Reasoning Attention (MRA) modules are introduced to integrate features along the temporal dimension to keep the video spatio-temporal consistent. Extensive experiments conducted on three widely-used public benchmarks, UCF101, HMDB51, and InfAR, indicate that the excellence of our approach over other state-of-the-art models in wild scenarios.
| Original language | English |
|---|---|
| Pages (from-to) | 2597-2601 |
| Number of pages | 5 |
| Journal | IEEE Signal Processing Letters |
| Volume | 29 |
| DOIs | |
| State | Published - 2022 |
| Externally published | Yes |
Keywords
- Action recognition
- semantic correspondence
- spatio-temporal alignment
Fingerprint
Dive into the research topics of 'Actor-Aware Alignment Network for Action Recognition'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver