Skip to main navigation Skip to search Skip to main content

CNN-GCN Coordinated Multimodal Frequency Network for Hyperspectral Image and LiDAR Classification

  • Haibin Wu
  • , Haoran Lv
  • , Aili Wang*
  • , Siqi Yan
  • , Gabor Molnar
  • , Liang Yu
  • , Minhui Wang
  • *Corresponding author for this work
  • Harbin University of Science and Technology
  • Julius Kühn Institute - Federal Research Centre for Cultivated Plants
  • Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Highlights: What are the main findings? We propose a novel CNN-GCN framework coordinated with wavelet transform for HSI and LiDAR classification. Its core innovation is a set of dedicated modules that work in concert to effectively balance local detail extraction with global contextual modeling. The proposed method achieves state-of-the-art classification performance, significantly outperforming existing advanced methods across three standard benchmark datasets. What is the implication of the main finding? The study provides an effective solution to key challenges in multimodal remote sensing, such as balancing local details with global contexts and enabling computationally efficient deep feature interaction. The framework’s superior generalization capability across diverse scenes demonstrates its strong potential as a reliable tool for enhancing accuracy in practical applications like environmental monitoring and urban planning. The existing multimodal image classification methods often suffer from several key limitations: difficulty in effectively balancing local detail and global topological relationships in hyperspectral image (HSI) feature extraction; insufficient multi-scale characterization of terrain features from light detection and ranging (LiDAR) elevation data; and neglect of deep inter-modal interactions in traditional fusion methods, often accompanied by high computational complexity. To address these issues, this paper proposes a comprehensive deep learning framework combining convolutional neural network (CNN), a graph convolutional network (GCN), and wavelet transform for the joint classification of HSI and LiDAR data, including several novel components: a Spectral Graph Mixer Block (SGMB), where a CNN branch captures fine-grained spectral–spatial features by multi-scale convolutions, while a parallel GCN branch models long-range contextual features through an enhanced gated graph network. This dual-path design enables simultaneous extraction of local detail and global topological features from HSI data; a Spatial Coordinate Block (SCB) to enhance spatial awareness and improve the perception of object contours and distribution patterns; a Multi-Scale Elevation Feature Extraction Block (MSFE) for capturing terrain representations across varying scales; and a Bidirectional Frequency Attention Encoder (BiFAE) to enable efficient and deep interaction between multimodal features. These modules are intricately designed to work in concert, forming a cohesive end-to-end framework, which not only achieves a more effective balance between local details and global contexts but also enables deep yet computationally efficient interaction across features, significantly strengthening the discriminability and robustness of the learned representation. To evaluate the proposed method, we conducted experiments on three multimodal remote sensing datasets: Houston2013, Augsburg, and Trento. Quantitative results demonstrate that our framework outperforms state-of-the-art methods, achieving OA values of 98.93%, 88.05%, and 99.59% on the respective datasets.

Original languageEnglish
Article number216
JournalRemote Sensing
Volume18
Issue number2
DOIs
StatePublished - Jan 2026
Externally publishedYes

Keywords

  • convolutional neural network (CNN)
  • graph convolutional network (GCN)
  • hyperspectral image (HSI)
  • light detection and ranging (LiDAR)
  • multimodal image classification

Fingerprint

Dive into the research topics of 'CNN-GCN Coordinated Multimodal Frequency Network for Hyperspectral Image and LiDAR Classification'. Together they form a unique fingerprint.

Cite this