Linear composition
Composing affine layers without activation is still only one affine map.
The source of neural-network expressiveness
Without nonlinear activations, stacked layers collapse into one linear transformation. Sigmoid, tanh, and ReLU break that collapse in different ways.
The essential idea: depth adds expressive power only when nonlinear transformations appear between linear layers.
Composing affine layers without activation is still only one affine map.
A hinge at zero allows many neurons to assemble piecewise-linear curves.
Switch activation functions and inspect both the output and derivative across the input range.
Nonlinear activations prevent linear collapse and balance expressiveness with trainable gradients.