Skip to main navigation Skip to search Skip to main content

Cross-modality fusion with local and global semantic feature correlation for weakly distinguishable object detection in aerial visible and infrared images

  • Maozhen Liu
  • , Xiaoguang Di*
  • , Ximing Li
  • , Zihao Wang
  • *Corresponding author for this work
  • Harbin Institute of Technology
  • Harbin

Research output: Contribution to journalArticlepeer-review

Abstract

To address the issue of missed detections caused by the similarity between weakly distinguishable objects and the background in both visible and infrared images, we utilize the correlation between local and global semantic features to find the difference clues between objects and the background. Specifically, we propose a novel Cross-Modality Fusion with Local and Global Semantic Feature Correlation Object Detection Network (CFLGNet), which effectively enhances the distinguishability. Firstly, we design a Semantic-Attention Mamba Fusion Block called SAMFB, which maps the cross-modal features into a hidden state space for interaction, and combines local and global semantic feature correlation to enhance the distinguishing features between the object and the background. SAMFB contains two branches: The Local Semantic Feature Correlation Module (LSCM) obtains an effective representation of local fine semantic feature through bilinear matrix computation, and this feature representation is embedded into another Global Semantic Feature Correlation Module with mamba (GSCM). GSCM captures the irregular spatial correlation through the scanning strategy of visual mamba, and then the local semantic feature extracted by the LSCM module are used as weights to suppress the spatially adjacent features with different representations. In addition, we propose a Lcal Feature Enhancement module based on Topological Invariants(LFETI) to reduce the problem of missed detection caused by incomplete object information, such as occlusion or truncation. Finally, the semantic feature correlation is combined with the original feature to enhance the distinguishing feature between the object and the background. Extensive experiments on the VEDAI and DroneVehicle benchmark datasets demonstrate that the proposed CFLGNet exhibits remarkable performance.

Original languageEnglish
Article number109404
JournalNeural Networks
Volume205
DOIs
StatePublished - Jan 2027

Keywords

  • Aerial images
  • Cross-modality fusion
  • Semantic feature correlation
  • Weakly distinguishable object detection

Fingerprint

Dive into the research topics of 'Cross-modality fusion with local and global semantic feature correlation for weakly distinguishable object detection in aerial visible and infrared images'. Together they form a unique fingerprint.

Cite this