Module II asked how we infer and update an uncertain state. Module III asks what a forecast distribution can support, and how trust changes near rare events, altered baselines, and unfamiliar conditions.
Module III · Exploration 9Predicting a Changing Planet
A Forecast Is a Distribution
The honest answer to “What will happen?” is often a structured set of possibilities.
A probabilistic forecast does not select one future. It describes how plausibility is distributed across several possible futures, and makes a claim that can be tested only across many comparable cases.
A single best estimate hides uncertainty in the current state, future forcing, model parameters, and model structure. An ensemble makes some of that uncertainty visible by generating several plausible evolutions.
Thin paths show plausible futures; the band contains the central 80%.
Five-day accumulated rainfallsynthetic conceptual forecast · millimetersDistribution at Day 5each mark is one ensemble member
14 of 40 members exceed the thresholdForecast probability: 35%
This outcome may fall near the center or in the tail. One outcome does not verify the distribution. Verification requires repeated forecast–outcome pairs.
Challenge
The ensemble is biased. Would 20, 100, or 1,000 members repair it?
Leave bias or the missing pathway unchanged, then compare the same forecast design at three ensemble sizes.
More members reduce sampling noise. They do not repair model bias or restore possibilities omitted from the ensemble design.
Two different questions about forecast quality
01 · Calibration
Do the probabilities mean what they say?
Among many comparable cases assigned a 70% probability, the event should occur about 70% of the time.
100 forecasts at 70%→70 events30 non-events
forecast probability ≈ observed frequency
Calibration asks whether stated probabilities agree with what happens across repeated comparable cases.
02 · Resolution / informativeness
Can the forecast distinguish situations with genuinely different risk?
A forecast has useful resolution when it gives different probabilities in situations whose event frequencies really are different.
Here “resolution” is a forecast-verification term. It does not mean spatial resolution.
Calibration gives probabilities meaning. Resolution makes them informative.
About sharpness. For continuous forecasts, sharpness describes how concentrated a predictive distribution is. A narrow distribution is not automatically good—it must also remain calibrated.
2 · Verify
Test the forecasting system across many cases
Each verification case contains a forecast probability pᵢ for a defined event and an observed outcome oᵢ, where oᵢ = 1 if the event occurs and oᵢ = 0 otherwise. One case cannot tell us whether a probability system is trustworthy.
Repeat the forecast many times:
(p₁, o₁), (p₂, o₂), …, (pN, oN)
The diagnostics below ask different questions about these same forecast–outcome pairs.
Changing sample size changes diagnostic noise, not the underlying forecasting system.
Choose a forecasting system
Does a 70% forecast behave like 70%?
Group cases with similar forecast probabilities. Within each group, calculate the mean issued probability p̄ and observed frequency ō, then plot (p̄, ō).
On the diagonal: probability and frequency agree.
Above: events happened more often than forecast.
Below: events happened less often than forecast.
Small samples produce noisy estimates, especially for probabilities issued rarely.
Reliability diagram
Inspect individual cases
Case 1: forecast 13%; event did not occur; Brier penalty 0.017.
Challenge
Both systems can be calibrated. Which one is more useful?
The same 100 realized outcomes are used for both systems: 40 lower-risk cases (10%), 40 baseline cases (30%), and 20 higher-risk cases (70%). Their expected total event frequency is 30%.
Left · Reliability
System A reliabilityBoth systems can be calibrated.
System A produces one point near (30%, 30%). System B produces three points near (10%, 10%), (30%, 30%), and (70%, 70%).
Right · Can the forecast separate risk?
Lower riskactual 10%issued 30%
Baseline riskactual 30%issued 30%
Higher riskactual 70%issued 30%
Brier score on the same cases: System A 0.214 · System B 0.166. Lower is better.
If both systems are calibrated, why is System B more useful?
calibrated ≠ informative The ideal forecast does both: its probability statements are calibrated, and they distinguish situations whose risks genuinely differ.
What does the score reward?
(0.20 − 1)²0.640
The Brier score gives a smooth squared penalty.
Proper scores reward honest probabilities, but different scores emphasize different mistakes.
Advanced note
For continuous outcomes, scores such as the continuous ranked probability score compare the full predictive distribution with the observed value.
3 · Decide
A warning is evaluated in a population of cases
Verification told us whether probabilities behave honestly. Decision problems ask something different: how should uncertain information be used when events are rare and mistakes have different consequences?
Event base rate · b1.0%
How often does the event occur before using this warning system?
Detection rate · d80%
Among event cases, what fraction receive a warning? Also called sensitivity or hit rate.
False-alarm rate · f5%
Among non-event cases, what fraction incorrectly receive a warning—not the fraction of warnings that are false.
Of all events, how many did we detect?80 / 100 = 80%
Of all warnings, how many were followed by the event?80 / 575 ≈ 14%
These percentages answer different conditional questions.
Advanced warning-system controls
A warning system can detect most rare events and still produce many false alarms because the non-event population is much larger. Usefulness depends on the base rate, warning-system performance, and consequences of misses and false alarms.
What if the event becomes more common?
Forecast ≠ decision
The same probability can support different actions
Suppose the forecast probability itself is already trusted. We now face a separate question: should we act?
Action A · Protect
Protection has cost C. Assume it avoids event loss completely.
Expected cost = CAction B · Do not protect
If the event occurs, loss is L. With forecast probability p:
Expected loss = pL
Protect now · C10Do not protect · pL20.0
Protect when C < pL, equivalently p > C/L.
Forecast probability 20%Action threshold 10%ACT under this toy cost–loss rule
The same scientifically valid 20% probability can rationally support different actions because the actions have different costs and consequences. forecast probability ≠ decision
The forecast supplies p. The decision threshold C/L comes from consequences.
What this toy decision model leaves out
It assumes one event, two actions, known costs and losses, full protection, no side effects, no uncertainty in cost estimates, and no competing users or equity considerations.
Its purpose is not to prescribe a real decision. It shows mathematically why a probability alone cannot determine an action.
Compare
Not every probabilistic prediction asks the same question
Weather forecasts, seasonal outlooks, and climate projections condition on different information, operate over different horizons, and should be interpreted and evaluated accordingly.
Where might this storm be three days from now?
Conditioned on
Estimated present atmosphere
Main uncertainty
Initial state + subsequent model evolution
Horizon
Days
Typical output
Trajectory ensemble / event probability
Evaluation
Many near-term forecast–outcome pairs arrive quickly
Match the question to the prediction
Where might a storm be on Friday?
Which temperature category is favored next season?
How may mid-century statistics differ under a scenario?
Communicate
The same distribution can tell several stories
Every summary preserves some information and hides something else. Choose the form that matches the question without pretending the discarded uncertainty has vanished.
Day 5 synthetic rainfall35% above 52 mm
Threshold probability. Useful when a specific consequence begins beyond a threshold.
A summary is a projection of the distribution, not a replacement for it.
A probability statement should specify
the event or quantity
the location and time period
the threshold, if one is used
the information on which the forecast is conditioned
the relevant baseline
performance in comparable past cases
important uncertainties not represented
Three probability traps
What sounds plausible, but is not enough?
01One outcome proves a probability forecast right or wrong.
A probability is a statement about repeated comparable cases. A well-calibrated 70% event should fail roughly 30% of the time.
02Calibration automatically makes a forecast useful.
Always predicting the climatological frequency can be calibrated and nearly uninformative. A useful forecast should add information without sacrificing calibration.
03The forecast probability determines the action.
A forecast describes uncertainty. A decision also includes consequences, costs, values, resources, and risk tolerance.
What should survive this experiment?
A useful forecast earns trust across many cases
A forecast is a distribution of plausible outcomes, not merely one preferred trajectory.
One outcome cannot verify a probability.
Calibration makes probability statements testable.
Calibration without informativeness may add little beyond the baseline.
Forecast quality and decision quality are connected but not identical.
Here, we treated a consequential event as a threshold and asked whether its probability was trustworthy. The next exploration moves into the tail itself: why some events become rare, amplified, or extreme, and why those mechanisms should not be confused.
Read the experiment carefully
What this Forecast Lab leaves out
Every forecast, outcome, regime, and warning case on this page is synthetic. The rainfall ensemble is a conceptual five-day system, not an operational weather forecast. The verification cases are independent and use stable statistical relationships unless the background-shift control is activated.
Finite samples make reliability estimates noisy, especially within subgroups.
More ensemble members reduce sampling noise but cannot repair bias or missing pathways.
Scores answer defined questions; no single score ranks every forecast for every use.
Real decisions include resources, communication, equity, and consequences not represented by a simple cost–loss ratio.
Sources, methods, and synthetic-data note
T. Gneiting and A. E. Raftery, “Strictly Proper Scoring Rules, Prediction, and Estimation,” 2007.
A. H. Murphy, “A New Vector Partition of the Probability Score,” 1973.
I. T. Jolliffe and D. B. Stephenson, Forecast Verification: A Practitioner’s Guide in Atmospheric Science, 2012.
D. S. Wilks, Statistical Methods in the Atmospheric Sciences, fourth edition, 2019.
Random experiments use deterministic pseudo-random seeds so forecast systems can be compared on the same cases. Ensemble probabilities are direct member fractions. Reliability uncertainty bars are approximate sampling intervals. The warning grid reports expected category counts from the selected rates.
Continue Module III
From trustworthy probabilities into the tail
A threshold forecast tells us how often an event may occur. The next exploration asks why rare or extreme events arise, and why an excursion, nonlinear amplification, and a change of regime are not the same mechanism.