Skip to main navigation Skip to search Skip to main content

SFOFusion: a task-oriented meta-learning framework for spatial–frequency fusion of infrared and visible images

  • Tao Jiang
  • , Hongyang Zhao*
  • , Yun Liu
  • , Jiayi Sun
  • , Xingdong Li
  • , Jing Jin
  • *Corresponding author for this work
  • College of Mechanical and Electrical Engineering, Northeast Forestry University
  • School of Astronautics, Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Infrared and visible image fusion plays a crucial role in various applications such as computer vision and remote sensing, especially in low-visibility environments, where it combines the powerful perception capability of infrared images in low light with the rich detail information of visible images. However, existing methods typically process image features separately in either the spatial or frequency domain, neglecting the deep fusion and optimization of cross-domain information. Moreover, few methods consider how to jointly optimize downstream tasks with spatial-frequency fusion models. To address these issues, we propose a task-oriented meta-learning framework, SFOFusion, which integrates phase information from the frequency domain with an explicit difference compensation mechanism in the spatial domain. The frequency domain part enhances cross-domain fusion accuracy by progressively merging amplitude and phase information, while the spatial domain retains high-frequency information through explicit compensation strategies. The training strategy uses an internal–external bi-level meta-learning scheme with epoch-level dynamic balancing. The internal–external loops introduce downstream task feedback into the fusion-loss optimization process, while the epoch-level balancing mechanism adaptively adjusts the relative importance of fusion fidelity and task performance to prevent either objective from dominating training. Experimental results demonstrate that SFOFusion significantly improves fusion performance on multiple datasets while maintaining model lightweight. It achieves a 4.7% improvement in mAP@0.5 for object detection and a 6.3% improvement in mIoU for semantic segmentation. Furthermore, SFOFusion reduces computational demands, with 25% fewer parameters and 30% less computation compared to traditional methods, highlighting its efficiency and practicality. The code is available at: https://github.com/Simon-JT/SFOFusion.

Original languageEnglish
Article number245401
JournalMeasurement Science and Technology
Volume37
Issue number24
DOIs
StatePublished - Jun 2026
Externally publishedYes

Keywords

  • frequency domain
  • image fusion
  • infrared and visible image
  • lightweight network
  • real-time image fusion

Fingerprint

Dive into the research topics of 'SFOFusion: a task-oriented meta-learning framework for spatial–frequency fusion of infrared and visible images'. Together they form a unique fingerprint.

Cite this