Abstract
To address the issue of missed detections caused by the similarity between weakly distinguishable objects and the background in both visible and infrared images, we utilize the correlation between local and global semantic features to find the difference clues between objects and the background. Specifically, we propose a novel Cross-Modality Fusion with Local and Global Semantic Feature Correlation Object Detection Network (CFLGNet), which effectively enhances the distinguishability. Firstly, we design a Semantic-Attention Mamba Fusion Block called SAMFB, which maps the cross-modal features into a hidden state space for interaction, and combines local and global semantic feature correlation to enhance the distinguishing features between the object and the background. SAMFB contains two branches: The Local Semantic Feature Correlation Module (LSCM) obtains an effective representation of local fine semantic feature through bilinear matrix computation, and this feature representation is embedded into another Global Semantic Feature Correlation Module with mamba (GSCM). GSCM captures the irregular spatial correlation through the scanning strategy of visual mamba, and then the local semantic feature extracted by the LSCM module are used as weights to suppress the spatially adjacent features with different representations. In addition, we propose a Lcal Feature Enhancement module based on Topological Invariants(LFETI) to reduce the problem of missed detection caused by incomplete object information, such as occlusion or truncation. Finally, the semantic feature correlation is combined with the original feature to enhance the distinguishing feature between the object and the background. Extensive experiments on the VEDAI and DroneVehicle benchmark datasets demonstrate that the proposed CFLGNet exhibits remarkable performance.
| Original language | English |
|---|---|
| Article number | 109404 |
| Journal | Neural Networks |
| Volume | 205 |
| DOIs | |
| State | Published - Jan 2027 |
Keywords
- Aerial images
- Cross-modality fusion
- Semantic feature correlation
- Weakly distinguishable object detection
Fingerprint
Dive into the research topics of 'Cross-modality fusion with local and global semantic feature correlation for weakly distinguishable object detection in aerial visible and infrared images'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver