What is backpropagation doing, and what does a vanishing gradient look like in practice?
Backpropagation applies the chain rule backwards through a network, reusing each layer's cached activations to get every parameter's gradient in one pass. Gradients vanish when many small factors multiply, and it shows as a stalled loss with early-layer weights barely moving while the last layers still train.