MPE StudioMath of Planet Earth

Module III · Exploration 11

Can AI Learn the Earth?

First, watch a machine learn one small relationship.

Modern forecast systems begin with observations, turn them into a reconstructed analysis, and use that state to predict what comes next. GraphCast showed how powerful this learned step can be.

observations→reconstructed analysis→AI prediction
Train a simple learner ↓
Thumbnail for Can AI Understand the Earth—or Just Predict It?

Watch first · short video

Can AI Understand the Earth—or Just Predict It?

Ask what an AI system learns from Earth data before testing where its predictions remain trustworthy.

Watch video ↗Applied Mathematics in Geosciences · Episode 11
Related: Why Simulations Must Respect Physics ↗
Then explore it yourself ↓
01 · Input → output

Start with the simplest possible learning problem

For this teaching experiment, x and y are generic variables. They are not meant to represent a specific atmospheric or oceanic quantity. The goal is to isolate the mathematics of learning.

y = f∗(x) = 0.8x + 0.35 sin(2x)−2 ≤ x ≤ 2
Reference relationship and 81 training examplesClick any archive point
Input xOutput ytraining range
Inputxi = 0.52
→
Targetyi = 0.722

The learner receives xi, produces a prediction ŷi, and compares that prediction with the target yi.

02 · The small neural network

A nonlinear function built from simple pieces

Follow one selected input through four hidden units. Every displayed value is calculated from the current weights.

ŷ = ∑j=14 vj tanh(wjx + bj) + c
Inputx
→→
Outputŷ
Forward pass for x = 0.52
z1 = w1x + b1 = -0.800
h1 = tanh(z1) = -0.664
v1h1-0.080v2h2-0.008v3h30.007v4h4-0.039bias c0.050
Prediction ŷ = -0.070
Prediction, target, and error
prediction-0.070target0.722error-0.793
ℓ = ½(ŷ − y)2 = 0.314
Click a chain-rule path
∂ℓ/∂ŷ = ŷ − y = -0.793
∂ℓ/∂v1 = (ŷ − y)h1 = 0.526
∂h1/∂z1 = 1 − h12 = 0.559
∂ℓ/∂w1 = (ŷ − y)v1(1 − h12)x = -0.028
∂ℓ/∂b1 = (ŷ − y)v1(1 − h12) = -0.053
Learning rate η
θnew = θold − η∇θL

Backpropagation is the chain rule organized as an algorithm. It tells every parameter how a small change would change the loss.

03 · Train on the archive

Now repeat the update across all examples

One gradient step changes the curve only a little. Repeated epochs make the learned function approach the reference relationship across the archive.

L(θ) = (1/N) ∑i=1N ½[fθ(xi) − yi]2
Target function and learned functionblue = target · orange = learner
Input xOutput ytraining range
Loss history0 epochs
training epoch
Archive MSE1.0183Parameters13Examples81
04 · Change the loss

The loss changes what the learner prioritizes

Keep the archive and architecture fixed. Change only the mistakes that count most.

LMSE = (1/N)∑ ½(ŷi−yi)2

Treat every archive example equally.

Central fitbestArchive tailsgoodRepeated useuntested
The “best” model depends on which errors the objective makes expensive.
05 · Test the learned function

Interpolation is not extrapolation

Inside the archive, neighboring examples constrain the learned curve. Outside it, many continuations can fit the same training data.

Input xOutput ytraining range
Interpolation

x = 0.65

Reference y = 0.857
Prediction ŷ = 0.838
Absolute error = 0.019

This input is bracketed by training examples.
Optional: why two familiar inputs can form an unfamiliar combination
Input a is in range ✓Input b is in range ✓But the pair (a,b) can lie outside the joint training cloud.
06 · Use the model repeatedly

A good one-step prediction is not automatically a good rollout

In closed loop, each prediction becomes the next input. Small local errors can be carried forward and reshaped by the map.

Reference trajectory and learned predictionsblue = reference · orange = prediction
use number t
ŷt+1 = fθ(ŷt)
et+1 ≈ f∗′(xt)et + bt

f∗′(xt)et carries earlier error forward. bt is the new local approximation error introduced at this step.

07 · Can AI learn the Earth?

What does the evidence actually support?

Choose the strongest conclusion justified by each result.

Scenario 1 of 4 · Familiar holdout

A model has low error on withheld examples drawn from the same input range.

08 · Back to Earth

The same logic scales to real scientific models

The toy network is deliberately small. The questions it reveals remain essential at Earth-system scale.

Application

Weather prediction

Skillful learned forecasts are real evidence for the variables, initialization, lead times, regimes, and evaluation data that were tested.

Application

Hybrid Earth-system models

A learned component changes the states it later receives after coupling, so the complete coupled model must be evaluated in closed loop.

Application

Climate response

Short-range weather skill does not by itself establish long-term response under new forcing, feedbacks, mean states, or slow components.

What to carry forward

A model learns the task we specify—and earns trust through the tests it survives.

training examples

→

forward propagation

→

loss

→

backpropagation

→

gradient descent

→

learned function

The scientific question is not only whether AI predicts well, but what relationship it learned, from which evidence, for which target, and within which domain.
Sources, methods, and synthetic-data note
  • Lam et al., “Learning skillful medium-range global weather forecasting,” Science, 2023.
  • Literature on distribution shift, shortcut learning, neural weather prediction, hybrid models, and closed-loop model stability.

The browser experiment uses deterministic synthetic teaching data and a small four-unit tanh learner. It is designed to expose changes in the scientific question, not to reproduce GraphCast or another operational forecasting system.

Continue the Lab

From a learned model to the whole model ladder