Abstract
Highlights: What are the main findings? We propose a novel CNN-GCN framework coordinated with wavelet transform for HSI and LiDAR classification. Its core innovation is a set of dedicated modules that work in concert to effectively balance local detail extraction with global contextual modeling. The proposed method achieves state-of-the-art classification performance, significantly outperforming existing advanced methods across three standard benchmark datasets. What is the implication of the main finding? The study provides an effective solution to key challenges in multimodal remote sensing, such as balancing local details with global contexts and enabling computationally efficient deep feature interaction. The framework’s superior generalization capability across diverse scenes demonstrates its strong potential as a reliable tool for enhancing accuracy in practical applications like environmental monitoring and urban planning. The existing multimodal image classification methods often suffer from several key limitations: difficulty in effectively balancing local detail and global topological relationships in hyperspectral image (HSI) feature extraction; insufficient multi-scale characterization of terrain features from light detection and ranging (LiDAR) elevation data; and neglect of deep inter-modal interactions in traditional fusion methods, often accompanied by high computational complexity. To address these issues, this paper proposes a comprehensive deep learning framework combining convolutional neural network (CNN), a graph convolutional network (GCN), and wavelet transform for the joint classification of HSI and LiDAR data, including several novel components: a Spectral Graph Mixer Block (SGMB), where a CNN branch captures fine-grained spectral–spatial features by multi-scale convolutions, while a parallel GCN branch models long-range contextual features through an enhanced gated graph network. This dual-path design enables simultaneous extraction of local detail and global topological features from HSI data; a Spatial Coordinate Block (SCB) to enhance spatial awareness and improve the perception of object contours and distribution patterns; a Multi-Scale Elevation Feature Extraction Block (MSFE) for capturing terrain representations across varying scales; and a Bidirectional Frequency Attention Encoder (BiFAE) to enable efficient and deep interaction between multimodal features. These modules are intricately designed to work in concert, forming a cohesive end-to-end framework, which not only achieves a more effective balance between local details and global contexts but also enables deep yet computationally efficient interaction across features, significantly strengthening the discriminability and robustness of the learned representation. To evaluate the proposed method, we conducted experiments on three multimodal remote sensing datasets: Houston2013, Augsburg, and Trento. Quantitative results demonstrate that our framework outperforms state-of-the-art methods, achieving OA values of 98.93%, 88.05%, and 99.59% on the respective datasets.
| Original language | English |
|---|---|
| Article number | 216 |
| Journal | Remote Sensing |
| Volume | 18 |
| Issue number | 2 |
| DOIs | |
| State | Published - Jan 2026 |
| Externally published | Yes |
Keywords
- convolutional neural network (CNN)
- graph convolutional network (GCN)
- hyperspectral image (HSI)
- light detection and ranging (LiDAR)
- multimodal image classification
Fingerprint
Dive into the research topics of 'CNN-GCN Coordinated Multimodal Frequency Network for Hyperspectral Image and LiDAR Classification'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver