Just Weather Scoring: Efficient End-to-End Nowcasting with Distributional Diffusion

Jannik Wiese*, Johannes Schusterbauer*, Tommaso Martorella, Björn Ommer
CompVis @ LMU Munich, Munich Center for Machine Learning (MCML)
*equal contribution
Comparison of CasCast, FREUD, and JWS: JWS generates radar forecasts with a single end-to-end model, eliminating separate compression and decoding stages. CRPS versus training TFLOPs for JWS, CasCast, and FREUD, including their separate training stages.

JWS simplifies probabilistic precipitation nowcasting with a single, end-to-end model that generates forecasts directly in data space, achieving competitive forecasting at substantially lower training cost.

TL;DR

  • Just Weather Scoring (JWS) generates probabilistic precipitation nowcasts directly in radar space with a single, end-to-end diffusion model.
  • Masked Asynchronous Diffusion (MAD) preserves clean observations while adapting noise levels to high-dimensional radar sequences.
  • The simple distributional diffusion model (sDDM) objective combines CRPS and squared error for few-step forecasting with strong probabilistic skill and further gains from more inference steps.

Abstract

Generative diffusion models are well-suited for probabilistic precipitation nowcasting, but existing approaches often rely on separately trained compression or deterministic forecasting components and remain costly at inference due to iterative denoising. We introduce Just Weather Scoring (JWS), a single-stage, end-to-end diffusion model which addresses both issues by forecasting directly in radar space and enabling few-step generation. Radar-space modeling greatly simplifies training and inference and eliminates uncertainty arising from lossy compression. JWS combines Masked Asynchronous Diffusion, a timestep-sampling scheme that preserves clean context while adapting diffusion training to high-dimensional spatio-temporal data, with a simple scoring-rule objective that aligns training with probabilistic forecasting and unlocks few-step generation. On the SEVIR and MeteoNet benchmarks, JWS achieves state-of-the-art probabilistic forecasting performance at reduced training and inference cost. Even our smallest model remains competitive using substantially fewer parameters and more than 17× faster inference.

Method

FREUD ensembles reconstructions of latent forecasts; JWS samples forecast ensembles directly in radar space.

JWS generates radar forecasts with a single transformer, avoiding separately trained compression and deterministic forecasting stages. Sampling multiple futures while keeping observations fixed estimates predictive uncertainty directly in radar space.


Masked Asynchronous Diffusion:

Correlated framewise diffusion timesteps sampled around a shared center, with selected context frames kept clean.

MAD keeps context frames clean, aligning training with forecasting from observed radar data. For remaining frames, it samples correlated noise levels around a shared center, adapted to spatial and temporal resolution. Per-frame noise levels and flexible context selection enable asynchronous denoising, varying confidence in observations or forecast priors, and filling gaps in partially observed sequences.


simple Distributional Diffusion:

Distributional diffusion represents conditional denoising distributions, enabling larger sampling steps than mean-velocity prediction.

sDDM linearly interpolates between pixelwise CRPS at high noise and mean squared error near clean data. This transitions from distributional supervision to regression, enabling few-step forecasting while retaining improvements from additional denoising steps.

Key Results

SEVIR nowcasting:

MethodStagesParamsCRPS ↓SSIM ↑HSS ↑CSI ↑RI* ↓
ConvLSTM114M0.02640.77490.52320.4102—
PredRNN147M0.02710.74970.51920.4045—
PhyDNet114M0.02530.76490.53110.4198—
SimVP116M0.02590.77720.52800.4153—
EarthFormer19M0.02510.77560.54110.4310—
NowcastNet135M0.02830.56960.53650.4152—
PreDiff†2105M0.02020.76480.49140.3875—
CasCast3402M0.02020.77970.56020.44010.3124
FlowCast2160M0.0182—0.58630.46510.6835
FREUD2521M0.01900.78410.50110.38640.1355
JWS-T/3219M0.01840.80160.51320.39480.0949
JWS-S/32124M0.01780.80570.53270.41160.0966
JWS-B/32194M0.01760.80920.54080.41800.0956
JWS-L/321340M0.01750.81220.54280.41830.1908
↳ with CFG = 1.51340M0.01770.80230.58940.45810.1815

Bold: best score. † Trained at 128 × 128. * RI computed with our evaluation pipeline.

Show MeteoNet resultsHide MeteoNet results
MethodStagesParamsCRPS ↓SSIM ↑HSS ↑CSI ↑RI* ↓
EarthFormer19M0.0224——0.2831—
NowcastNet135M0.0277——0.2955—
PreDiff2105M0.0197——0.2546—
CasCast3402M0.0180——0.3156—
FREUD2521M0.01930.73120.20820.1417—
JWS-T/3219M0.01430.78870.40260.28510.2376
JWS-S/32124M0.01410.79170.40790.28910.3100
JWS-B/32194M0.01440.79670.42060.29550.3858
JWS-L/321340M0.01380.79740.41860.29590.0773
↳ with CFG = 1.51340M0.01410.79040.48160.34480.5719

Even the 9M-parameter JWS-T outperforms prior methods in CRPS. * RI computed with our evaluation pipeline.

JWS achieves the best CRPS and SSIM on SEVIR and MeteoNet with a single-stage model. Classifier-free guidance further improves HSS and CSI.


Fast probabilistic forecasts:

CRPS versus time in seconds per forecast, comparing JWS-B sampling budgets with CasCast and FREUD.
CRPS versus forecast latency. Bubble size indicates parameter count.

One-step JWS-B generates a 10-member ensemble in less than a second, with CRPS competitive with CasCast and FREUD. More denoising steps further improve forecast quality.


Calibrated uncertainty:

SEVIR rank histograms: JWS is closer to the uniform ideal than CasCast and FREUD. Reliability index over forecast lead time: JWS maintains lower RI than CasCast and FREUD.
Rank histograms (left) and reliability index over lead time (right). Uniform ranks and lower RI are better.

JWS produces flatter rank histograms than CasCast and FREUD, indicating more reliable uncertainty estimates. This calibration advantage persists across forecast lead times.


Inference time scaling:

Left: JWS-B CRPS and CSI across ensemble sizes and function evaluations (NFE). Right: JWS-S compared with Flow Matching, Shortcut, and FGN across ensemble sizes and NFE.
Ensemble and NFE scaling for JWS-B (left), and comparisons with FM, Shortcut, and FGN (right).

JWS benefits from larger ensembles and more denoising steps. With just one function evaluation, JWS-S outperforms Flow Matching, Shortcut models, and FGN in matched comparisons.

Qualitative Forecasts

+5 min

Ours

GT Forecast

FREUD

GT Forecast

CasCast

GT Forecast

FlowCast

GT Forecast
GT Forecast
GT Forecast
GT Forecast
GT Forecast
GT Forecast
GT Forecast
GT Forecast
GT Forecast

Citation

@misc{wiese2026jws,
    title = {Just Weather Scoring: Efficient End-to-End Nowcasting with Distributional Diffusion},
    author = {Wiese, Jannik and Schusterbauer, Johannes and Martorella, Tommaso and Ommer, Bj{\"o}rn},
    year = {2026}
}