Skip to main navigation Skip to search Skip to main content

Retain, Blend, and Exchange: A Quality-Aware Spatial-Stereo Fusion Approach for Event Stream Recognition

  • Lan Chen
  • , Dong Li
  • , Xiao Wang*
  • , Pengpeng Shao
  • , Wei Zhang
  • , Yaowei Wang
  • , Yonghong Tian
  • , Jin Tang
  • *Corresponding author for this work
  • Anhui University
  • School of Computer Science and Technology, Anhui University
  • Tsinghua University
  • Peng Cheng Laboratory
  • Harbin Institute of Technology Shenzhen
  • Peking University

Research output: Contribution to journalArticlepeer-review

Abstract

Current event stream-based pattern recognition models typically present the event stream as the point cloud, voxel, image, and the like, and formulate multiple deep neural networks to acquire their features. Although considerable results can be achieved in simple cases, however, the performance of the model might be restricted by monotonous modality expressions, sub-optimal fusion, and readout mechanisms. In this article, we put forward a novel dual-stream framework for event stream-based pattern recognition through differentiated fusion, which is called EFV++. It models two common event representations simultaneously, i.e., event images and event voxels. The spatial and three-dimensional stereo information can be separately learned by making use of Transformer and Graph Neural Network (GNN). We believe the features of each representation still contain both efficient and redundant features and a sub-optimal solution may be obtained if we directly fuse them without differentiation. Thus, we divide each feature into three levels and retain high-quality features, blend medium-quality features, and exchange low-quality features. The enhanced dual features will be provided to the fusion Transformer together with bottleneck features. In addition, we introduce a novel hybrid interaction readout mechanism to enhance the diversity of features as final representations. Comprehensive experiments validate that the framework we have proposed attains cutting-edge performance on a variety of extensively utilized event stream-based classification datasets. Particularly, we have realized a freshly pioneering performance on the Bullying10 k dataset, precisely 90.51%, and this outpaces the runner-up by +2.21%.

Original languageEnglish
Pages (from-to)8926-8939
Number of pages14
JournalIEEE Transactions on Multimedia
Volume27
DOIs
StatePublished - 12 Nov 2025
Externally publishedYes

Keywords

  • Event camera
  • event stream-based classification
  • graph neural networks
  • multi-view representation
  • transformer

Fingerprint

Dive into the research topics of 'Retain, Blend, and Exchange: A Quality-Aware Spatial-Stereo Fusion Approach for Event Stream Recognition'. Together they form a unique fingerprint.

Cite this