Abstract
Accurate solar forecasting is essential for grid integration but remains challenging due to the stochastic nature of clouds. Despite over a decade of research on machine learning to augment physical forecasting frameworks, progress has been incremental. Recent advances in foundation models have reignited optimism for a step change in solar forecasting. We propose VISIF (Vision-Integrated Solar Irradiance Forecaster), a method that leverages pre-trained large vision–language models for satellite-based solar irradiance forecasting. Rather than assuming that natural image–language alignment transfers directly to satellite remote sensing and ground radiometer data, VISIF adapts a pre-trained large vision–langauge model (LVLM) through task-specific interfaces: a modified multi-spectral visual embedding layer, a learnable time-series tokenizer, and a forecasting decoder. The LVLM backbone then serves as a pre-trained multimodal sequence processor for fusing advected satellite imagery with historical ground observations. Experiments across geographically diverse stations show that VISIF consistently outperforms state-of-the-art unimodal and multimodal baselines, reducing the mean absolute error by more than 21% relative to the CrossViViT benchmark. Scaling analysis further indicates that small and mid-sized backbones generalize better across climates. Code is provided for reproducibility.
| Original language | English |
|---|---|
| Article number | 100864 |
| Journal | Atmospheric and Oceanic Science Letters |
| DOIs | |
| State | Accepted/In press - 2026 |
| Externally published | Yes |
Keywords
- Large vision-language models
- Remote sensing
- Solar irradiance forecasting
Fingerprint
Dive into the research topics of 'Solar forecasting with large vision–language models'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver