MPE StudioMath of Planet Earth
Module III begins

Module II asked how we infer and update an uncertain state. Module III asks what a forecast distribution can support, and how trust changes near rare events, altered baselines, and unfamiliar conditions.

Module III · Exploration 9Predicting a Changing Planet

A Forecast Is a Distribution

The honest answer to “What will happen?” is often a structured set of possibilities.

A probabilistic forecast does not select one future. It describes how plausibility is distributed across several possible futures, and makes a claim that can be tested only across many comparable cases.

Enter the Forecast Lab ↓
consequential rainfall thresholdsame initialized stateENSEMBLE EVENT PROBABILITY5 / 17 = 29%initializationforecast lead
One initialized state produces an ensemble of plausible rainfall futures. The fraction crossing a stated threshold becomes an event probability.
Playing
Thumbnail for From Chaos to Probability

Watch first · short video

From Chaos to Probability

See why a forecast distribution can be more honest and more useful than a single predicted future.

Watch video ↗Applied Mathematics in Geosciences · Episode 9
Related: Why Prediction Has Limits (Even with Perfect Models) ↗
Then explore it yourself ↓
Regional outlook

Tomorrow

30%

chance of rain

Before the lab

It rains the next day. Was the forecast wrong?

1 · Build

One forecast, many possible futures

A single best estimate hides uncertainty in the current state, future forcing, model parameters, and model structure. An ensemble makes some of that uncertainty visible by generating several plausible evolutions.

Thin paths show plausible futures; the band contains the central 80%.

Five-day accumulated rainfallsynthetic conceptual forecast · millimeters
0306090120thresholdDay 1Day 2Day 3Day 4Day 5
Distribution at Day 5each mark is one ensemble member
0306090120
14 of 40 members exceed the thresholdForecast probability: 35%

This outcome may fall near the center or in the tail. One outcome does not verify the distribution. Verification requires repeated forecast–outcome pairs.

Challenge

The ensemble is biased. Would 20, 100, or 1,000 members repair it?

Leave bias or the missing pathway unchanged, then compare the same forecast design at three ensemble sizes.

More members reduce sampling noise. They do not repair model bias or restore possibilities omitted from the ensemble design.
Two different questions about forecast quality
01 · Calibration

Do the probabilities mean what they say?

Among many comparable cases assigned a 70% probability, the event should occur about 70% of the time.

100 forecasts at 70%→70 events30 non-events

forecast probability ≈ observed frequency

Calibration asks whether stated probabilities agree with what happens across repeated comparable cases.

02 · Resolution / informativeness

Can the forecast distinguish situations with genuinely different risk?

lower-risk cases 10%ordinary-risk cases 30%higher-risk cases 70%

A forecast has useful resolution when it gives different probabilities in situations whose event frequencies really are different.

Here “resolution” is a forecast-verification term. It does not mean spatial resolution.
Calibration gives probabilities meaning. Resolution makes them informative.

About sharpness. For continuous forecasts, sharpness describes how concentrated a predictive distribution is. A narrow distribution is not automatically good—it must also remain calibrated.

2 · Verify

Test the forecasting system across many cases

Each verification case contains a forecast probability pᵢ for a defined event and an observed outcome oᵢ, where oᵢ = 1 if the event occurs and oᵢ = 0 otherwise. One case cannot tell us whether a probability system is trustworthy.

Repeat the forecast many times:

(p₁, o₁), (p₂, o₂), …, (pN, oN)

The diagnostics below ask different questions about these same forecast–outcome pairs.

Number of verification cases
Advanced / sampling demonstration

Changing sample size changes diagnostic noise, not the underlying forecasting system.

Choose a forecasting system

Does a 70% forecast behave like 70%?

Group cases with similar forecast probabilities. Within each group, calculate the mean issued probability p̄ and observed frequency ō, then plot (p̄, ō).

  • On the diagonal: probability and frequency agree.
  • Above: events happened more often than forecast.
  • Below: events happened less often than forecast.

Small samples produce noisy estimates, especially for probabilities issued rarely.

Reliability diagram
00252550507575100100forecast probability (%)observed frequency (%)

Inspect individual cases

Case 1: forecast 13%; event did not occur; Brier penalty 0.017.

Challenge

Both systems can be calibrated. Which one is more useful?

The same 100 realized outcomes are used for both systems: 40 lower-risk cases (10%), 40 baseline cases (30%), and 20 higher-risk cases (70%). Their expected total event frequency is 30%.

Left · Reliability

System A reliability
00252550507575100100forecast probability (%)observed frequency (%)
Both systems can be calibrated.

System A produces one point near (30%, 30%). System B produces three points near (10%, 10%), (30%, 30%), and (70%, 70%).

Right · Can the forecast separate risk?

Lower riskactual 10%issued 30%
Baseline riskactual 30%issued 30%
Higher riskactual 70%issued 30%

Brier score on the same cases: System A 0.214 · System B 0.166. Lower is better.

If both systems are calibrated, why is System B more useful?

calibrated ≠ informative
The ideal forecast does both: its probability statements are calibrated, and they distinguish situations whose risks genuinely differ.
What does the score reward?
Outcome
Score
(0.20 − 1)²0.640

The Brier score gives a smooth squared penalty.

Proper scores reward honest probabilities, but different scores emphasize different mistakes.

Advanced note

For continuous outcomes, scores such as the continuous ranked probability score compare the full predictive distribution with the observed value.

3 · Decide

A warning is evaluated in a population of cases

Verification told us whether probabilities behave honestly. Decision problems ask something different: how should uncertain information be used when events are rare and mistakes have different consequences?

Event base rate · b1.0%

How often does the event occur before using this warning system?

Detection rate · d80%

Among event cases, what fraction receive a warning? Also called sensitivity or hit rate.

False-alarm rate · f5%

Among non-event cases, what fraction incorrectly receive a warning—not the fraction of warnings that are false.

very rarecommon
EventNo eventWarning
Correct warning · Nbd80
False alarm · N(1−b)f495
No warning
Missed event · Nb(1−d)20
Correct quiet case · N(1−b)(1−f)9,405

Of all events, how many did we detect?80 / 100 = 80%

Of all warnings, how many were followed by the event?80 / 575 ≈ 14%

These percentages answer different conditional questions.

Advanced warning-system controls
Number of cases
A warning system can detect most rare events and still produce many false alarms because the non-event population is much larger.
Usefulness depends on the base rate, warning-system performance, and consequences of misses and false alarms.

What if the event becomes more common?

Forecast ≠ decision

The same probability can support different actions

Suppose the forecast probability itself is already trusted. We now face a separate question: should we act?

Action A · Protect

Protection has cost C. Assume it avoids event loss completely.

Expected cost = C
Action B · Do not protect

If the event occurs, loss is L. With forecast probability p:

Expected loss = pL
Protect now · C10
Do not protect · pL20.0

Protect when C < pL, equivalently p > C/L.

Forecast probability 20%Action threshold 10%ACT under this toy cost–loss rule
The same scientifically valid 20% probability can rationally support different actions because the actions have different costs and consequences.
forecast probability ≠ decision

The forecast supplies p. The decision threshold C/L comes from consequences.

What this toy decision model leaves out

It assumes one event, two actions, known costs and losses, full protection, no side effects, no uncertainty in cost estimates, and no competing users or equity considerations.

Its purpose is not to prescribe a real decision. It shows mathematically why a probability alone cannot determine an action.

Compare

Not every probabilistic prediction asks the same question

Weather forecasts, seasonal outlooks, and climate projections condition on different information, operate over different horizons, and should be interpreted and evaluated accordingly.

Where might this storm be three days from now?

Conditioned on
Estimated present atmosphere
Main uncertainty
Initial state + subsequent model evolution
Horizon
Days
Typical output
Trajectory ensemble / event probability
Evaluation
Many near-term forecast–outcome pairs arrive quickly

Match the question to the prediction

Where might a storm be on Friday?

Which temperature category is favored next season?

How may mid-century statistics differ under a scenario?

Communicate

The same distribution can tell several stories

Every summary preserves some information and hides something else. Choose the form that matches the question without pretending the discarded uncertainty has vanished.

Day 5 synthetic rainfall35% above 52 mm

Threshold probability. Useful when a specific consequence begins beyond a threshold.

A summary is a projection of the distribution, not a replacement for it.

A probability statement should specify

  • the event or quantity
  • the location and time period
  • the threshold, if one is used
  • the information on which the forecast is conditioned
  • the relevant baseline
  • performance in comparable past cases
  • important uncertainties not represented
Three probability traps

What sounds plausible, but is not enough?

01One outcome proves a probability forecast right or wrong.

A probability is a statement about repeated comparable cases. A well-calibrated 70% event should fail roughly 30% of the time.

02Calibration automatically makes a forecast useful.

Always predicting the climatological frequency can be calibrated and nearly uninformative. A useful forecast should add information without sacrificing calibration.

03The forecast probability determines the action.

A forecast describes uncertainty. A decision also includes consequences, costs, values, resources, and risk tolerance.

What should survive this experiment?

A useful forecast earns trust across many cases

  1. A forecast is a distribution of plausible outcomes, not merely one preferred trajectory.
  2. One outcome cannot verify a probability.
  3. Calibration makes probability statements testable.
  4. Calibration without informativeness may add little beyond the baseline.
  5. Forecast quality and decision quality are connected but not identical.

Here, we treated a consequential event as a threshold and asked whether its probability was trustworthy. The next exploration moves into the tail itself: why some events become rare, amplified, or extreme, and why those mechanisms should not be confused.

Read the experiment carefully

What this Forecast Lab leaves out

Every forecast, outcome, regime, and warning case on this page is synthetic. The rainfall ensemble is a conceptual five-day system, not an operational weather forecast. The verification cases are independent and use stable statistical relationships unless the background-shift control is activated.

  • Finite samples make reliability estimates noisy, especially within subgroups.
  • More ensemble members reduce sampling noise but cannot repair bias or missing pathways.
  • Scores answer defined questions; no single score ranks every forecast for every use.
  • Real decisions include resources, communication, equity, and consequences not represented by a simple cost–loss ratio.
Sources, methods, and synthetic-data note
  • T. Gneiting and A. E. Raftery, “Strictly Proper Scoring Rules, Prediction, and Estimation,” 2007.
  • A. H. Murphy, “A New Vector Partition of the Probability Score,” 1973.
  • I. T. Jolliffe and D. B. Stephenson, Forecast Verification: A Practitioner’s Guide in Atmospheric Science, 2012.
  • D. S. Wilks, Statistical Methods in the Atmospheric Sciences, fourth edition, 2019.

Random experiments use deterministic pseudo-random seeds so forecast systems can be compared on the same cases. Ensemble probabilities are direct member fractions. Reliability uncertainty bars are approximate sampling intervals. The warning grid reports expected category counts from the selected rates.

Continue Module III

From trustworthy probabilities into the tail

A threshold forecast tells us how often an event may occur. The next exploration asks why rare or extreme events arise, and why an excursion, nonlinear amplification, and a change of regime are not the same mechanism.