Derivatives on a computer: finite differences, the chain rule, backpropagation
Part 1 defined the derivative as the limit of the secant slope,
and ran into a problem: from function values alone a computer only ever gets an estimate. This part measures how good that estimate can be, and then shows what neural networks do instead — differentiate the formula once by rules, compose the rules with the chain rule, and run the result backwards through a computation graph. That last step is backpropagation, where the first lecture of Andrej Karpathy’s Neural Networks: Zero to Hero is heading. Every panel below is live: change something and the numbers are recomputed, never typed in.
The curve is the same as in part 1:
The derivative as a function
Since every point has its own slope, the derivative is itself a function, . It is rarely computed from the limit directly. A handful of rules, each proven once from the definition, is enough:
- a sum differentiates term by term: ;
- a constant factor stays: ;
- powers: ;
- a product: ;
- ;
- — a function that is its own slope (it and its multiples are the only ones);
- the chain rule for a composition: .
Together they reach , the activation of the first neuron in the lecture: the rules give . That slope is at and fades towards as grows, which is why a neuron driven deep into saturation barely learns — whatever gradient arrives at it is multiplied by almost nothing.
For our curve, the chain rule turns into , and the rest is term by term:
That formula is what the dashed tangents in both parts are drawn from.
Where , the tangent is horizontal. On the plotted range that happens 7 times; where changes from negative to positive the curve has a local minimum — the deepest at , where — and where it changes from positive to negative, a local maximum.
Not every function has a derivative everywhere. has slope to the left of and to the right; the secant has no single limit at , so is not differentiable there. The activation function ReLU, , has the same corner, and neural-network libraries simply agree to use there.
Zoom into the corner and see for yourself. The panel draws two secants from the point, one to the left and one to the right, with tied to the zoom. On a smooth curve, zooming in straightens the picture and the two secants fold into one line. At the corner of the picture never changes: at every zoom it is the same V, the left secant stays at and the right at . There is no single number for the derivative to be.
The corner of |x| under magnification: the left and right secants from x = 0 keep slopes −1 and +1 at every zoom, while on the smooth x² they merge into one line; ReLU shows the 0 that libraries use at its corner.
How small should h be on a computer?
In exact arithmetic, a smaller is always better. On a computer it is not.
Part 1’s secant slope , the forward difference, subtracts two nearly equal numbers. A 64-bit float carries about 16 significant digits, so the difference loses as many of them as and share. The rounding error that is left is then divided by the tiny .
Two errors pull in opposite directions:
- truncation error — the curve is not a line — is about and shrinks with ;
- rounding error is about , where is the float precision, and grows as shrinks.
Their sum is smallest near .
The central difference uses points on both sides. The terms cancel, its truncation error is of order , and its best moves up to about .
The panel computes both formulas in ordinary 64-bit arithmetic and plots their error against ; drag across it to move .
Log-log plot of the error of the forward and central differences of f at x = 3 against h: both errors fall as h shrinks until rounding takes over and they rise again, the forward difference bottoming out near h = 1e-8 and the central one near 1e-5.
Read the plot from right to left, the way shrinks. At first both lines fall with slopes 1 and 2: every tenfold decrease of divides the forward error by 10 and the central error by 100. Then rounding takes over and both lines turn up into noise.
At the forward difference is most accurate at , with error . The central difference reaches at — about 450 times more accurate, with 1000 times larger.
At both estimates are exactly : the float nearest to is itself, so .This is why numerical derivatives are used to check analytic ones — a “gradient check” with the central difference and around — and not to train networks: each estimate costs extra function evaluations per input, and the answer is good to only about ten digits.
The chain rule on a computation graph
Real formulas are compositions of simple operations. Break
into steps: , , . This is a computation graph.
The forward pass computes values left to right. With : .
The backward pass answers a different question: for every node, how much does change when that node is nudged? That is the derivative , called its gradient. A node’s value and its gradient are different numbers.
Each operation only needs to know its own local derivative:
- for : and ;
- for : — addition passes the gradient through unchanged;
- for : and — multiplication hands each input the value of the other.
Rates along a chain multiply. The lecture borrows George F. Simmons’s example: if a car travels twice as fast as a bicycle, and the bicycle four times as fast as a walking man, the car travels times as fast as the man. The chain rule multiplies local derivatives along the path in exactly that way:
Move the backward-pass slider: every notch processes one operation, from right to left, and labels each edge it crosses with its local derivative. Then drag any input sideways on the graph, or move its slider.
The computation graph of L = (a·b + c)·s with each node's value and gradient; a backward-pass slider processes one operation at a time from the output towards the inputs, labelling edges with local derivatives, and dragging an input compares the change in L with the gradient's prediction.
It starts from : nudging the output by changes it by . At the end , and dragging checks them the old way. Drag from to and moves by — exactly . The gradient is that prediction: how much the output moves per unit nudge of the input. Here it is exact, because is linear in each input on its own; for a curved function it holds only for small nudges, which is the whole content of .
Now move all four inputs at once. “Step along the gradient” adds to every input times its own gradient, and rises from to . The gradient predicted a rise of ; the true one is . The gap is the products of inputs — , — that change together and that a first-order prediction cannot see: is linear in each input alone, not in all of them at once. Stepping against the gradient would lower instead, and that is the whole of training, as the last section shows.
A ”—” before a node is reached means “not computed yet”, which is not the same as zero. Drag to and every gradient to its left becomes zero, although none of those nodes has value zero: when the last factor is zero, nothing upstream can move .
This is backpropagation. One backward sweep yields the gradient of the single output with respect to every node, at a cost of a small constant times the forward pass. A network has one loss and millions of parameters, and that ratio is why it trains backwards and not forwards.
One variable, several paths: gradients add
What if a value is used more than once? Take
The variable enters the graph three times: twice as an input of and once in . Each use is a separate path to , and each path sends back its own contribution: from each side of the product, from the second term. The multivariable chain rule says the derivative is their sum:
That is why a backpropagation engine writes grad += ... and not grad = .... The second button replays the bug: each contribution overwrites the one before, only the last survives, and the tangent drawn from that gradient misses the curve.
The graph of L = x·x + k·x, where x reaches L along three paths and sends back three contributions; next to it the curve L(x) with a tangent drawn from the accumulated gradient, which misses the curve when contributions overwrite instead of add.
Drag the point to with . The sum is zero and the tangent is horizontal, yet no path has disappeared: . Contributions can cancel, and a zero gradient does not mean nothing depends on .
Two different things are called “accumulating” gradients, and it pays not to mix them up. Within one backward pass, contributions must add. Between training steps, the stored gradients must be reset to zero first — otherwise the previous step’s gradient is added to the new one. The first is the chain rule; forgetting the second is a bug.
Order matters too. A node may pass its gradient on only after every use of it has reported back, which is why the backward pass visits nodes in reverse topological order.
Back to the walk downhill
Part 1 started from gradient descent, , on a curve of one variable. Training a neural network is that same loop: becomes millions of weights, becomes the loss on the training data, becomes the gradient vector, and backpropagation computes it in one sweep. “Step along the gradient” on the graph above is one step of that loop, run uphill; stepping against the gradient runs it downhill. The rest of the lecture builds exactly that on top of the tiny engine in these panels.
Short answers
Does every function have a derivative?
No. |x| has a corner at 0: the slope is −1 on the left and +1 on the right, so the secant has no single limit. ReLU, max(0, x), has the same corner; neural-network libraries simply pick a value there, usually 0.
What is the difference between a derivative and a gradient?
A derivative is the rate of change of a function of one variable. A gradient collects the partial derivatives of a function of many variables, one per input, into a vector; it points in the direction of steepest increase.
Numerical, symbolic or automatic differentiation — which is which?
Numerical differentiation estimates the slope from function values at nearby points and trades truncation error for rounding error. Symbolic differentiation manipulates formulas. Automatic differentiation, which backpropagation is an instance of, applies the chain rule to the actual sequence of operations a program executed and gives exact derivatives up to floating-point rounding.
Why does backpropagation run backwards?
A network has one loss and millions of parameters. Going backwards from the loss, one sweep over the graph yields the derivative of that single output with respect to every input, at a cost of a small constant times the forward pass. Going forwards would take one sweep per parameter.
Why are gradients accumulated with += in backpropagation?
Because a value that is used in several places influences the output along several paths, and by the multivariable chain rule its derivative is the sum of the contributions of all those paths. Assigning instead of adding keeps only the last path.
Further reading
- A. Karpathy, “The spelled-out intro to neural networks and backpropagation: building micrograd”, Neural Networks: Zero to Hero, lecture 1, 2022. The lecture both parts follow.
- M. Spivak, Calculus, 4th ed., Publish or Perish, 2008. Chapters 9–10: the definition of the derivative and the rules, proven properly.
- J. Nocedal, S. J. Wright, Numerical Optimization, 2nd ed., Springer, 2006. Chapter 8: finite differences, their errors, and automatic differentiation.
- A. G. Baydin, B. A. Pearlmutter, A. A. Radul, J. M. Siskind, “Automatic differentiation in machine learning: a survey”, Journal of Machine Learning Research 18 (2018), 1–43.
- A. Griewank, A. Walther, Evaluating Derivatives: Principles and Techniques of Algorithmic Differentiation, 2nd ed., SIAM, 2008. The reference on reverse-mode differentiation.