Abstract
Multi-modal crowd counting is an essential yet challenging task that uses rich information to enhance the accuracy of crowd counting in complex environments. In this paper, we place particular emphasis on modal asymmetry in the fusion process of multi-modal crowd counting and propose a novel and powerful asymmetric modal fusion approach to effectively utilize the different modal information for multi-modal fusion. In addition, we propose a self-supervised enhanced training scheme, termed the Modal Swapping Fusion Consistency (MSFC) mechanism, in which the fusion features dominated by one modality can be emulated by the fusion features dominated by another modality to ensure consistency among different fusion features for further refining the fusion process. Extensive experiments conducted on multiple multi-modal crowd counting datasets, including RGB-thermal and RGB-depth, demonstrate that our approach significantly outperforms existing methods. Our project will be available at: https://github.com/Mr-Monday/AMFA.
| Original language | English |
|---|---|
| Article number | 112768 |
| Journal | Pattern Recognition |
| Volume | 172 |
| DOIs | |
| State | Published - Apr 2026 |
| Externally published | Yes |
Keywords
- Multi-modal crowd counting
- Multi-modal fusion
- Prompt learning
- Self-supervised learning
Fingerprint
Dive into the research topics of 'Asymmetric modal fusion for multi-modal crowd counting'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver