Skip to main navigation Skip to search Skip to main content

空—地多视角行为识别的判别信息增量学习方法

Translated title of the contribution: Discriminative information incremental learning for Air-ground multi-view action recognition
  • Liu Wenxuan
  • , Zhong Xian*
  • , Xu Xiaoyu
  • , Zhou Zhuo
  • , Jiang Kui
  • , Wang Zheng
  • , Bai Xiang
  • *Corresponding author for this work
  • Wuhan University of Technology
  • State Key Laboratory of Maritime Technology and Safety
  • Wuhan University
  • School of Computer Science and Technology, Harbin Institute of Technology
  • Huazhong University of Science and Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Objective With the increasing demand for urban security of people, ground devices are combined with air devices, such as drones, for identifying action in air-ground scenarios. Meanwhile, the extensive ground-based camera networks and a wealth of ground surveillance data can offer reliable support to these aerial surveillance devices. How to effectively utilize the mobility of these aerial devices is a topic that warrants further research. Existing multi-view action recognition methods focus only on the difference in discriminative action information when the horizontal spatial view changes, but do not consider the difference in discriminative action information when the vertical spatial view changes. The high mobility of aerial perspectives can lead to changes in the vertical spatial perspective. According to the principles of perspective, observing the same object from different heights results in a significant change in appearance. This, in turn, causes substantial differences in the appearance of the same person’s actions when observed from high-altitude and ground-level perspectives. These significant variations in action appearance are referred to as differences in discriminative action information, and they pose a challenge for traditional multi-view action recognition methods in effectively addressing the issue of vertical spatial perspective changes. Method When the viewing perspective aligns with the objects being observed in the same horizontal spatial plane, the most comprehensive and rich discriminative action information can be observed. Networks can easily learn and comprehend this information. However, when the viewing perspective is in a different horizontal spatial plane from the observed objects, inclined perspective occurs, resulting in a significant change in action appearance. This transition from a ground-level perspective to an aerial perspective leads to insufficiently observed information and a reduction in the discriminative action information. When networks attempt to learn and understand this information, misclassifications are more likely to occur. Therefore, on the basis of the amount of discriminative action information, ground-level perspective information can be considered as easily learned and understood simple information, while aerial perspective information can be seen as complex information that is challenging to learn and understand. In fact, the human brain follows a progressive learning process when dealing with various types of information, prioritizing the processing of simple information and using the learned simple information to assist in learning complex information. In the vertical spatial multi-view action recognition task, differences in perspectives and environmental influences lead to varying amounts of discriminative action information observed at different heights. In this chapter, we adopt a brain-like approach. We rank samples from the aerial perspective on the basis of the amount of discriminative action information they contain. Complex samples contain less discriminative action information, and networks find them challenging to learn and understand. Simple samples contain more discriminative action information and are easier for networks to learn and comprehend. We then distill discriminative action information separately from simple and complex samples. Within the same action category, despite differences in the amount of discriminative action information between simple and complex samples, the represented action categories should have commonalities. Therefore, by using the discriminative action information incremental learning method, we incrementally inject the rich discriminative action information learned from simple samples into the feature information of complex samples. This approach addresses the issue of complex samples carrying insufficient discriminative action information, allowing complex samples to learn more discriminative action information with the assistance of simple samples. Thus, networks can learn and understand complex samples easily. This paper proposes a discriminative action information incremental learning (DAIL) for multi-view action recognition in complex air-ground scenes and to distinguish the ground view from the air view on the basis of the view height and the amount of information. This paper utilizes a neuromorphic learning knowledge referred to as“ordered incremental progression”to distill discriminative action information for different views separately. Discriminative action information is incremented from the ground-view (simple) samples into the air-view (complex) samples to assist the network in learning and understanding the air-view samples. Result The method is experimentally validated on two datasets, namely, Drone-Action and unmanned aerial vehicle (UAV). The accuracy of the two datasets is improved by 18. 0% and 16. 2%, respectively, compared with that of the current state-of-the-art method SBP. Compared with the strong baseline method, our method reduces the parameters by 2. 4 M and the FLOPS by 6. 9 G on the UAV dataset. To validate the effectiveness of our proposed method in scenarios involving both ground-level and aerial perspectives, we introduced two datasets:N-UCLA (comprising samples exclusively from ground-based cameras with rich discriminative behavior information) and Drone-Action (comprising a mix of ground-level and aerial samples, where aerial samples contain relatively limited discriminative behavior information). A joint analysis of discriminative behavior information ranking was conducted on these datasets. Our findings indicate that enhancing complex samples using simpler ones significantly improves the network’s feature learning capacity. Conversely, attempting the reverse can lead reduce the accuracy. This observation aligns with the way the human brain processes information, embodying the concept of progressive learning. Conclusion In this study, we proposed DAIL for multi-view action recognition in complex air-ground scenes and to distinguish the ground view from the air view on the basis of the view height and the amount of information. Experiment results show that our model outperforms several state-of-the-art multi-view approaches and improves the performance.

Translated title of the contributionDiscriminative information incremental learning for Air-ground multi-view action recognition
Original languageChinese (Traditional)
Pages (from-to)130-147
Number of pages18
JournalJournal of Image and Graphics
Volume30
Issue number1
DOIs
StatePublished - Jan 2025
Externally publishedYes

Fingerprint

Dive into the research topics of 'Discriminative information incremental learning for Air-ground multi-view action recognition'. Together they form a unique fingerprint.

Cite this