Gradient descent
Each parameter moves opposite the local loss gradient.
Forward prediction and backward credit assignment
Training alternates a forward pass that computes a prediction with a backward pass that differentiates the loss and updates every parameter.
The essential idea: backpropagation uses the chain rule to assign each weight its share of the prediction error.
Each parameter moves opposite the local loss gradient.
Backpropagation multiplies local derivatives along computational paths.
Follow Charlotte's sample through the forward pass, loss, gradient, and update, while the convergence curve records how loss changes with iteration count.
Neural networks learn by forward evaluation, backpropagated gradients, and repeated parameter updates.