MPE StudioMath of Planet Earth
Machine Learning Toolkit · 16

Nonlinear Activation Functions

The source of neural-network expressiveness

Without nonlinear activations, stacked layers collapse into one linear transformation. Sigmoid, tanh, and ReLU break that collapse in different ways.

watchone idea
→
manipulateone example
→
leave withone intuition

The essential idea: depth adds expressive power only when nonlinear transformations appear between linear layers.

Watch the concept

One Concept · One Example

Nonlinear Activation Functions video thumbnail▶

Nonlinear Activation Functions

Presented by Charlotte Moser

Watch on YouTube ↗

What to notice

The idea in 30 seconds

Break the linear collapse

Linear composition

Composing affine layers without activation is still only one affine map.

W2(W1x+b1)+b2=Wx+b

ReLU

A hinge at zero allows many neurons to assemble piecewise-linear curves.

ReLU(x)=max(0,x)
Explore

Compare activation shapes

Switch activation functions and inspect both the output and derivative across the input range.

KEY TAKEAWAY

Nonlinear activations prevent linear collapse and balance expressiveness with trainable gradients.