Why a probability is not forecast skill
A 60% probability of above-normal temperature says nothing about how often the model is right. What skill means for seasonal forecasts and why it is missing here.
Two different questions
A probability answers “what share of this model’s members fall in this category?”. Skill answers “when this model said 60%, how often was it right?”. The first is a property of one forecast; the second is a property of the system, measured over many past forecasts against observations. A probability map on this site tells you the first and nothing about the second.
Why the distinction matters
Seasonal forecast skill varies enormously by region, season, variable and lead. Tropical sea-surface temperature at short lead is often well forecast. Winter precipitation over northern Europe at lead 4 is close to climatology in most systems. A 65% upper-tercile probability in both places looks identical on the map but should be trusted very differently. Only a verification study can tell you which.
Reliability
A well-calibrated system that says 60% is right about 60% of the time. Raw ensemble probabilities from a single model are often overconfident, because members share the same model errors and spread less than reality does. Centres publish reliability diagrams and skill scores for their systems; the C3S site provides verification maps for each contributing model. Consult those before treating any probability here as a calibrated confidence.
What this site does and does not provide
This site derives probabilities directly from member counts, with empirical thresholds from the 1993–2016 hindcast, and applies no calibration. It does not yet compare forecasts with observations, and it presents no skill scores. Agreement between several models is a useful qualitative signal but is not a substitute for verification, and “neutral” on a tercile summary is not the same as a confident forecast of near-normal conditions.
Read probability maps as a description of what a model’s ensemble contains. Combine them with published skill assessments, with agreement across the five models here, and with consistency across runs before acting on them.