Derivatives on a computer: finite differences, the chain rule, backpropagation

Part 1 defined the derivative as the limit of the secant slope,

f′(x)=lim⁡h→0f(x+h)−f(x)h,f'(x) = \lim_{h \to 0} \frac{f(x+h) - f(x)}{h},

and ran into a problem: from function values alone a computer only ever gets an estimate. This part measures how good that estimate can be, and then shows what neural networks do instead — differentiate the formula once by rules, compose the rules with the chain rule, and run the result backwards through a computation graph. That last step is backpropagation, where the first lecture of Andrej Karpathy’s Neural Networks: Zero to Hero is heading. Every panel below is live: change something and the numbers are recomputed, never typed in.

The curve is the same as in part 1:

f(x)=3sin⁡(2.5x)+0.4x2−1.5x.f(x) = 3\sin(2.5x) + 0.4x^2 - 1.5x .

The derivative as a function

Since every point has its own slope, the derivative is itself a function, x↦f′(x)x \mapsto f'(x). It is rarely computed from the limit directly. A handful of rules, each proven once from the definition, is enough:

Together they reach tanh⁡x=e2x−1e2x+1\tanh x = \dfrac{e^{2x} - 1}{e^{2x} + 1}, the activation of the first neuron in the lecture: the rules give (tanh⁡x)′=1−tanh⁡2x(\tanh x)' = 1 - \tanh^2 x. That slope is 11 at 00 and fades towards 00 as ∣x∣|x| grows, which is why a neuron driven deep into saturation barely learns — whatever gradient arrives at it is multiplied by almost nothing.

For our curve, the chain rule turns sin⁡(2.5x)\sin(2.5x) into cos⁡(2.5x)⋅2.5\cos(2.5x) \cdot 2.5, and the rest is term by term:

f′(x)=7.5cos⁡(2.5x)+0.8x−1.5.f'(x) = 7.5\cos(2.5x) + 0.8x - 1.5 .

That formula is what the dashed tangents in both parts are drawn from.

Where f′(x)=0f'(x) = 0, the tangent is horizontal. On the plotted range that happens 7 times; where f′f' changes from negative to positive the curve has a local minimum — the deepest at x≈1.885x \approx 1.885, where f≈−4.406f \approx -4.406 — and where it changes from positive to negative, a local maximum.

Not every function has a derivative everywhere. ∣x∣|x| has slope −1-1 to the left of 00 and +1+1 to the right; the secant has no single limit at 00, so ∣x∣|x| is not differentiable there. The activation function ReLU, max⁡(0,x)\max(0, x), has the same corner, and neural-network libraries simply agree to use 00 there.

Zoom into the corner and see for yourself. The panel draws two secants from the point, one to the left and one to the right, with hh tied to the zoom. On a smooth curve, zooming in straightens the picture and the two secants fold into one line. At the corner of ∣x∣|x| the picture never changes: at every zoom it is the same V, the left secant stays at −1-1 and the right at +1+1. There is no single number for the derivative to be.

The corner of |x| under magnification: the left and right secants from x = 0 keep slopes −1 and +1 at every zoom, while on the smooth x² they merge into one line; ReLU shows the 0 that libraries use at its corner.

How small should h be on a computer?

In exact arithmetic, a smaller hh is always better. On a computer it is not.

Part 1’s secant slope (f(x+h)−f(x))/h\bigl(f(x+h) - f(x)\bigr)/h, the forward difference, subtracts two nearly equal numbers. A 64-bit float carries about 16 significant digits, so the difference loses as many of them as f(x+h)f(x+h) and f(x)f(x) share. The rounding error that is left is then divided by the tiny hh.

Two errors pull in opposite directions:

Their sum is smallest near h≈ε≈1.5⋅10−8h \approx \sqrt{\varepsilon} \approx 1.5 \cdot 10^{-8}.

The central difference (f(x+h)−f(x−h))/2h\bigl(f(x+h) - f(x-h)\bigr)/2h uses points on both sides. The f′′f'' terms cancel, its truncation error is of order h2h^2, and its best hh moves up to about ε3≈6.1⋅10−6\sqrt[3]{\varepsilon} \approx 6.1 \cdot 10^{-6}.

The panel computes both formulas in ordinary 64-bit arithmetic and plots their error against hh; drag across it to move hh.

Log-log plot of the error of the forward and central differences of f at x = 3 against h: both errors fall as h shrinks until rounding takes over and they rise again, the forward difference bottoming out near h = 1e-8 and the central one near 1e-5.

Read the plot from right to left, the way hh shrinks. At first both lines fall with slopes 1 and 2: every tenfold decrease of hh divides the forward error by 10 and the central error by 100. Then rounding takes over and both lines turn up into noise.

At x=3x = 3 the forward difference is most accurate at h≈4⋅10−9h \approx 4 \cdot 10^{-9}, with error 1.3⋅10−81.3 \cdot 10^{-8}. The central difference reaches 3⋅10−113 \cdot 10^{-11} at h≈4⋅10−6h \approx 4 \cdot 10^{-6} — about 450 times more accurate, with hh 1000 times larger.

At h=10−16h = 10^{-16} both estimates are exactly 00: the float nearest to 3+10−163 + 10^{-16} is 33 itself, so f(x+h)−f(x)=0f(x+h) - f(x) = 0.

This is why numerical derivatives are used to check analytic ones — a “gradient check” with the central difference and hh around 10−510^{-5} — and not to train networks: each estimate costs extra function evaluations per input, and the answer is good to only about ten digits.

The chain rule on a computation graph

Real formulas are compositions of simple operations. Break

L=(a⋅b+c)⋅sL = (a \cdot b + c) \cdot s

into steps: e=a⋅be = a \cdot b, d=e+cd = e + c, L=d⋅sL = d \cdot s. This is a computation graph.

The forward pass computes values left to right. With a=2, b=−3, c=10, s=−2a = 2,\ b = -3,\ c = 10,\ s = -2: e=−6, d=4, L=−8e = -6,\ d = 4,\ L = -8.

The backward pass answers a different question: for every node, how much does LL change when that node is nudged? That is the derivative ∂L/∂node\partial L / \partial \text{node}, called its gradient. A node’s value and its gradient are different numbers.

Each operation only needs to know its own local derivative:

Rates along a chain multiply. The lecture borrows George F. Simmons’s example: if a car travels twice as fast as a bicycle, and the bicycle four times as fast as a walking man, the car travels 2⋅4=82 \cdot 4 = 8 times as fast as the man. The chain rule multiplies local derivatives along the path in exactly that way:

∂L∂a=∂L∂d⋅∂d∂e⋅∂e∂a=s⋅1⋅b=(−2)⋅1⋅(−3)=6.\frac{\partial L}{\partial a} = \frac{\partial L}{\partial d} \cdot \frac{\partial d}{\partial e} \cdot \frac{\partial e}{\partial a} = s \cdot 1 \cdot b = (-2)\cdot 1 \cdot(-3) = 6 .

Move the backward-pass slider: every notch processes one operation, from right to left, and labels each edge it crosses with its local derivative. Then drag any input sideways on the graph, or move its slider.

The computation graph of L = (a·b + c)·s with each node's value and gradient; a backward-pass slider processes one operation at a time from the output towards the inputs, labelling edges with local derivatives, and dragging an input compares the change in L with the gradient's prediction.

It starts from ∂L/∂L=1\partial L / \partial L = 1: nudging the output by ε\varepsilon changes it by ε\varepsilon. At the end ∂L/∂a=6, ∂L/∂b=−4, ∂L/∂c=−2, ∂L/∂s=4\partial L/\partial a = 6,\ \partial L/\partial b = -4,\ \partial L/\partial c = -2,\ \partial L/\partial s = 4, and dragging checks them the old way. Drag aa from 22 to 2.52.5 and LL moves by 33 — exactly ∂L/∂a⋅Δa=6⋅0.5\partial L/\partial a \cdot \Delta a = 6 \cdot 0.5. The gradient is that prediction: how much the output moves per unit nudge of the input. Here it is exact, because LL is linear in each input on its own; for a curved function it holds only for small nudges, which is the whole content of f(x+h)≈f(x)+f′(x) hf(x+h) \approx f(x) + f'(x)\,h.

Now move all four inputs at once. “Step along the gradient” adds to every input 0.010.01 times its own gradient, and LL rises from −8.00-8.00 to −7.29-7.29. The gradient predicted a rise of 0.01⋅(62+42+22+42)=0.720.01 \cdot (6^2 + 4^2 + 2^2 + 4^2) = 0.72; the true one is 0.710.71. The gap is the products of inputs — a⋅ba \cdot b, e⋅se \cdot s — that change together and that a first-order prediction cannot see: LL is linear in each input alone, not in all of them at once. Stepping against the gradient would lower LL instead, and that is the whole of training, as the last section shows.

A ”—” before a node is reached means “not computed yet”, which is not the same as zero. Drag ss to 00 and every gradient to its left becomes zero, although none of those nodes has value zero: when the last factor is zero, nothing upstream can move LL.

This is backpropagation. One backward sweep yields the gradient of the single output with respect to every node, at a cost of a small constant times the forward pass. A network has one loss and millions of parameters, and that ratio is why it trains backwards and not forwards.

One variable, several paths: gradients add

What if a value is used more than once? Take

L=x⋅x+k⋅x.L = x \cdot x + k \cdot x .

The variable xx enters the graph three times: twice as an input of x⋅xx \cdot x and once in k⋅xk \cdot x. Each use is a separate path to LL, and each path sends back its own contribution: xx from each side of the product, kk from the second term. The multivariable chain rule says the derivative is their sum:

∂L∂x=x+x+k=2x+k.\frac{\partial L}{\partial x} = x + x + k = 2x + k .

That is why a backpropagation engine writes grad += ... and not grad = .... The second button replays the bug: each contribution overwrites the one before, only the last survives, and the tangent drawn from that gradient misses the curve.

The graph of L = x·x + k·x, where x reaches L along three paths and sends back three contributions; next to it the curve L(x) with a tangent drawn from the accumulated gradient, which misses the curve when contributions overwrite instead of add.

Drag the point to x=−1.5x = -1.5 with k=3k = 3. The sum is zero and the tangent is horizontal, yet no path has disappeared: −1.5−1.5+3=0-1.5 - 1.5 + 3 = 0. Contributions can cancel, and a zero gradient does not mean nothing depends on xx.

Two different things are called “accumulating” gradients, and it pays not to mix them up. Within one backward pass, contributions must add. Between training steps, the stored gradients must be reset to zero first — otherwise the previous step’s gradient is added to the new one. The first is the chain rule; forgetting the second is a bug.

Order matters too. A node may pass its gradient on only after every use of it has reported back, which is why the backward pass visits nodes in reverse topological order.

Back to the walk downhill

Part 1 started from gradient descent, x←x−η f′(x)x \leftarrow x - \eta\, f'(x), on a curve of one variable. Training a neural network is that same loop: xx becomes millions of weights, ff becomes the loss on the training data, f′(x)f'(x) becomes the gradient vector, and backpropagation computes it in one sweep. “Step along the gradient” on the graph above is one step of that loop, run uphill; stepping against the gradient runs it downhill. The rest of the lecture builds exactly that on top of the tiny engine in these panels.

Short answers

Does every function have a derivative?

No. |x| has a corner at 0: the slope is −1 on the left and +1 on the right, so the secant has no single limit. ReLU, max(0, x), has the same corner; neural-network libraries simply pick a value there, usually 0.

What is the difference between a derivative and a gradient?

A derivative is the rate of change of a function of one variable. A gradient collects the partial derivatives of a function of many variables, one per input, into a vector; it points in the direction of steepest increase.

Numerical, symbolic or automatic differentiation — which is which?

Numerical differentiation estimates the slope from function values at nearby points and trades truncation error for rounding error. Symbolic differentiation manipulates formulas. Automatic differentiation, which backpropagation is an instance of, applies the chain rule to the actual sequence of operations a program executed and gives exact derivatives up to floating-point rounding.

Why does backpropagation run backwards?

A network has one loss and millions of parameters. Going backwards from the loss, one sweep over the graph yields the derivative of that single output with respect to every input, at a cost of a small constant times the forward pass. Going forwards would take one sweep per parameter.

Why are gradients accumulated with += in backpropagation?

Because a value that is used in several places influences the output along several paths, and by the multivariable chain rule its derivative is the sum of the contributions of all those paths. Assigning instead of adding keeps only the last path.

Further reading