Positive covariance
Large values tend to occur with large values, and small values with small values.
How can two quantities vary together, and how can we measure that relationship without confusing scale, shape, or cause?
Many scientific questions involve more than one quantity. Temperature at two locations may rise and fall together. Wind and pressure may change in opposite directions. A model variable may appear closely related to an observation.
Covariance and correlation give us a language for describing these relationships. But they do not tell us exactly the same thing, and neither one by itself tells us why a relationship exists.
01 Center the data → 02 Measure joint variation → 03 Normalize and interpret
Start with paired observations (xi, yi). What matters is where a point lies relative to the average of each variable. Same-sign deviations make a positive centered product; opposite-sign deviations make a negative one.
Covariance begins by centering both variables. Each observation then contributes according to the product of its two deviations from the mean.
A single point does not determine covariance. It summarizes whether same-sign or opposite-sign deviations dominate across the dataset.
Move one point far from both means, then move it across a mean line. Watch its magnitude and sign change.
Covariance is the average product of two centered variables. It asks whether departures from the two means tend to have the same sign or opposite signs.
Estimating both means from the same sample uses one degree of freedom. Dividing by n−1 gives the usual unbiased sample covariance under standard sampling assumptions.
Large values tend to occur with large values, and small values with small values.
Large values tend to occur with small values of the other variable.
Positive and negative centered products largely cancel. This does not rule out a nonlinear relationship.
Covariance contains useful information, but its numerical value depends on the units used to measure the variables.
sx = 2.73 · sy = 2.28 · covariance = 6.12 · r = 0.98
sx* = 2.73 · covariance = 6.12 · r = 0.98
Changing an offset does not change covariance. Multiplying the scale changes covariance, even though the visible relationship has not fundamentally changed.
Pearson correlation is normalized covariance. It measures the direction and strength of a linear association.
Pearson correlation looks for a linear pattern. Nature is not required to organize itself along a straight line.
Pearson r = 0.99
Zero correlation is not the same as independence.
A Pearson correlation near zero says there is little linear association. A nonlinear dependence can still be strong.
Because covariance and correlation depend on distances from the mean, observations far from the center can contribute strongly.
A scatterplot can reveal co-variation. It cannot, by itself, identify the mechanism that produced it.
Observed X–Y correlation: 0.88
A large correlation is evidence of association, not a complete causal explanation. Direction, common drivers, selection effects, and other mechanisms require additional information or assumptions.
Earth-system models do not contain only two variables. The pairwise covariance idea extends naturally to a covariance matrix.
X₁ temperature-like · X₂ wind-like · X₃ ocean-memory-like
Selected entry C12: covariance 2.17
Foundation 04 · Principal Component Analysis
PCA uses covariance structure to identify directions along which a multivariable system varies most strongly. Coming next
Foundation 05 · The One-Dimensional Kalman Update
In data assimilation, covariance describes uncertainty and how information about one variable can update another. Coming next
Spatial coherenceMeasurements at nearby locations often vary together. Covariance helps describe how information is shared across space.
Coupled variablesAtmospheric, oceanic, and land variables can co-vary. Correlation is a first description; physical interpretation requires more.
UncertaintyModel errors and observational uncertainty can also co-vary. Covariance is central to uncertainty quantification, data assimilation, and prediction.
In Earth science, covariance is not merely a descriptive statistic. It is also a mathematical representation of how variability and uncertainty are organized across a system.
Covariance tells us how variations are connected. Correlation puts that connection on a common scale. The challenge is not only to calculate them, but to understand what they reveal, and what they do not.
Definitions and visual teaching choices follow standard sample covariance and Pearson-correlation treatments, including the NIST/SEMATECH e-Handbook of Statistical Methods. All datasets here are deterministic or reproducibly generated synthetic teaching data, not observational Earth-system records.