虎嗅

Why Did Backpropagation Change Neural Networks?

原文:为什么反向传播改变了神经网络?

Summary of Key Points

Backpropagation is the “unsung hero” behind training multi-layer neural networks. By utilizing the chain rule from calculus, it efficiently distributes the final error to each layer’s parameters, overcoming the challenge of training deep networks that was previously insurmountable. This mathematical technique, which emerged in the 1970s, was neglected for over a decade due to the AI winter before being widely recognized in 1986 with papers by Hinton and others. Today, it serves as the foundational mechanism in all deep learning frameworks (such as PyTorch and TensorFlow), powering modern AI applications ranging from image recognition to large models like GPT. Although it has some minor issues (such as gradient vanishing), it remains the most general and stable approach.

1. The “Deadlock” of the AI Winter: Why Couldn’t Multi-Layer Neural Networks Be Trained?

When the perceptron (a single-layer neural network) was introduced in 1958, the media hailed it as a breakthrough that would enable machines to “self-replicate” soon. However, in 1969, MIT’s Minsky demonstrated mathematically that a single-layer perceptron could not even solve a simple problem like XOR—the decision boundary required a curve, which a single-layer perceptron could only represent with a straight line. More critically, no one knew how to train multi-layer networks: the error from the output layer could be used to adjust the parameters of the final layer, but the intermediate “hidden layers” were like black boxes—no one knew which layer should bear the responsibility for the error (this is known as the “credit allocation problem”).

Old methods failed: random perturbation techniques were akin to guessing; increasing the number of parameters led to an explosion in computational complexity; layer-by-layer training only provided local insights and could not find the global optimum; expert systems relied on manually written rules, which proved too cumbersome. Without funding and interest in research, AI entered its first “winter,” lasting for 15 years.

2. The “Magic” of the Chain Rule: How Does Backpropagation Determine the Responsibility of Each Parameter?

The core of backpropagation is not new mathematics but the application of the chain rule, which we learn in middle school, to its extreme. Imagine a neural network as a series of functions nested within each other. The chain rule allows us to trace the error from the output layer back through the network, step by step, to determine how much each parameter contributed to the error (i.e., its gradient).

For example, if there is a problem with the final product in a factory, backpropagation would work backward, starting from the output layer to identify where the mistake occurred, then checking the previous layers, all the way back to the initial input data. The key is that this process is computationally inexpensive—similar to forward propagation (from input to output) and it reuses the intermediate results from forward propagation, avoiding unnecessary calculations.

3. The Neglected Genius: How Did Backpropagation Go From Obscurity to Center of Attention?

The idea of backpropagation existed long before its widespread adoption: in 1970, Finnish researcher Linnainmaa mentioned it in his master’s thesis, but no one connected it to neural networks. In 1974, Harvard doctoral student Werbos proposed using it for training multi-layer networks, yet his paper went unnoticed.

It wasn’t until 1986 that Hinton (later recognized as a pioneer of deep learning) and others published a paper in Nature demonstrating that backpropagation could indeed train effective multi-layer networks, enabling hidden layers to learn useful features (such as identifying edges in images). At that time, symbolicism (expert systems) was dominant, but Hinton’s team persisted, and eventually, backpropagation gained recognition throughout the field, helping AI emerge from the AI winter.

4. Today’s AI: Backpropagation—the Invisible “Mastermind”

When you use `loss.backward()` in PyTorch or `GradientTape` in TensorFlow, you are essentially triggering backpropagation. Its design principles have influenced all deep learning frameworks:

  • Local computation, global optimization: Each node calculates its own gradient, but together they lead to the global optimum.
  • Computational graphs: Frameworks record the forward propagation process, which is then used to propagate gradients backward.
  • Gradients as signals of responsibility: The gradient of each parameter indicates how much error it should be responsible for, guiding parameter adjustment.

From YOLO object detection to GPT models, from AlphaGo to AlphaFold, all modern AI models rely on backpropagation.

5. Backpropagation’s Challenges and Future Directions

Backpropagation is not without flaws:

  • Gradient vanishing/ explosion: Errors can become too small or too large during propagation; this issue has been largely addressed with techniques like the ReLU activation function and ResNet residual connections.
  • Biological plausibility: Neurons in the real brain do not propagate errors as precisely as backpropagation does; Hinton himself has explored alternative approaches that mimic the brain’s workings (e.g., the Forward-Forward algorithm proposed in 2022).
  • Hardware dependence: GPUs’ parallel computing power is essential for backpropagation’s matrix operations; without them, training large models would be impossible.

However, no alternative has yet surpassed backpropagation in terms of effectiveness, efficiency, and engineering feasibility. It remains the cornerstone of the AI field.

In conclusion: Truly great algorithms are not always obvious from the start; they emerge when the time is right to reveal their power to change the world. Backpropagation is such an algorithm.