Skip to main navigation Skip to search Skip to main content

Implicit Illumination-Aware Representation With Cross-Modal Prefusion Alignment for Universal Multispectral Pedestrian Detection

  • Yan Gong
  • , Lei Lin
  • , Yang Luo
  • , Hao Liu
  • , Yongsheng Gao*
  • , Jie Zhao
  • , Ziying Song
  • , Xiaoxi Hu
  • *Corresponding author for this work
  • Harbin Institute of Technology
  • X Division JD Logistics
  • Xiamen University
  • Beijing Jiaotong University

Research output: Contribution to journalArticlepeer-review

Abstract

Traditional pedestrian detection methods based on red-green-blue (RGB) images struggle in adverse illumination, but a key capability required for pedestrian detection is all-day detection due to its critical role in diverse applications, e.g., security, surveillance, and autonomous driving. To address this issue, multispectral pedestrian detection attempts to introduce thermal images to supplement the RGB images, since they can be captured based on heat radiation difference without relying on external light sources. However, how to fuse the two modalities effectively is still lacking in-depth investigation. To prompt this field, we propose an implicit illumination-aware representation to address the limited availability of specific illumination labels in existing multispectral datasets, coupled with a prefusion feature alignment strategy to reconcile spatial misalignments of identical objects across modalities. We also identify four critical fusion challenges, revealing persistent limitations in existing multispectral detectors’ ability to holistically address these issues, particularly regarding underdeveloped cross-modal interactions and suboptimal cross-domain feature fusion. To this end, we propose a universal multispectral pedestrian detection paradigm (UMPDP), which includes a modality alignment module (MAM) for adaptive feature space alignment, a differential modality fusion module (DMFM) to enhance the relationship of different modalities, and a task-conditioned illumination module (TCIM) to dynamically adjust network weights based on illumination condition. Extensive experiments on KAIST and CVC-14 datasets demonstrate the general effectiveness of our proposed method. Code is available at https://github.com/gongyan1/UMPDP

Original languageEnglish
Pages (from-to)3247-3261
Number of pages15
JournalIEEE Transactions on Neural Networks and Learning Systems
Volume37
Issue number7
DOIs
StatePublished - 1 Jul 2026

Keywords

  • Feature aligned and fusion
  • multimodal fusion
  • multispectral pedestrian detection
  • task-conditioned networks

Fingerprint

Dive into the research topics of 'Implicit Illumination-Aware Representation With Cross-Modal Prefusion Alignment for Universal Multispectral Pedestrian Detection'. Together they form a unique fingerprint.

Cite this