What is a derivative? Slope, limit, chain rule, interactively
The derivative of a function at a point is the number
In words: how fast the output changes per unit change of the input, measured so close to that nothing but the local behaviour of matters.
The same number has a second reading that is more useful than the first. For small ,
Near , the function behaves like a straight line with slope — the tangent. Everything on this page is that one line read in different directions: the tangent is its picture, finite differences estimate it, the chain rule composes it, backpropagation computes it for millions of inputs at once, and gradient descent uses it to walk downhill.
The route follows the first lecture of Andrej Karpathy’s Neural Networks: Zero to Hero, where the derivative is where building a neural network from scratch starts. Every panel below is live: change something and the numbers are recomputed, never typed in.
The lecture works with the parabola . This page uses a deliberately uneven curve instead, so its slope changes size and sign from place to place and there is more than one valley to fall into:
One point has no slope
Through a single point of a curve pass infinitely many straight lines. Which of them “has the slope of the curve”?
Drag the point along the curve and turn the line by its round handle. From a distance, plenty of lines look about right. The inset magnifies a small window around the point, and there the difference shows: a wrong line keeps crossing the curve at an angle however far you zoom, while one line merges with the curve. That line is the tangent, and its slope is .
A wavy curve with one marked point and a line through it whose slope the reader sets; an inset magnifies the point, where a wrong line crosses the curve at an angle and only the tangent merges with it.
The legend measures what the lens shows: the gap between curve and line divided by the distance from the point, at , and . At a flat line gives — the ratio closes in on and stays there, and every other wrong slope settles the same way on its own error. Turn the line to slope and the same three numbers become : ten times smaller for every tenfold zoom. Move the point, and the slope that worked a moment ago no longer does. A derivative is not one number for the whole curve but a new number at every point.
Why exactly one line? Expand the curve around . The gap between it and a line of slope , per unit of distance , is
As shrinks, everything but the first term disappears. Any error in the slope survives the zoom; only lets the ratio go to zero. This, and not “the line that touches the curve at one point”, is what a tangent is: a line can touch a curve once and still cross it, and a tangent can cross the curve (at an inflection point it does).
From secant to tangent: the limit
The point alone gives no slope, so take a second one. The points and fix a line, the secant, and its slope is the ratio of the legs of the orange triangle:
It is the tangent of the secant’s angle to the horizontal, — rise over run.
Now drag the second point towards the first. The slider does the same job: its right end is always the edge of the plot and its left end a single pixel, whatever the zoom. As shrinks, the secant turns onto the tangent.
The curve with a point x and a second point x + h joined by a secant; as h shrinks the secant rotates onto the dashed tangent, and zooming in shows the curve becoming indistinguishable from a straight line.
At the true slope is . The secant gets there step by step:
| step | secant slope | error |
|---|---|---|
| 1 | 6.6 | |
| 0.1 | 0.86 | |
| 0.01 | 0.084 | |
| 0.001 | 0.0084 |
At the secant even has the wrong sign: it jumps over the crest just past . Every further tenfold decrease of makes the error ten times smaller.
Now zoom in — with the zoom slider, or a pinch on a trackpad. The second point keeps its place on screen, so shrinks with the view, and the curve straightens under it: the secant turns onto the tangent without being touched. At a thousandth of the original width the curve is indistinguishable from its tangent. This is why the limit exists at all: a smooth function is, locally, linear. The derivative is the slope of that local line.
The derivative as a function
Since every point has its own slope, the derivative is itself a function, . It is rarely computed from the limit directly. A handful of rules, each proven once from the definition, is enough:
- a sum differentiates term by term: ;
- a constant factor stays: ;
- powers: ;
- a product: ;
- ;
- — a function that is its own slope (it and its multiples are the only ones);
- the chain rule for a composition: .
Together they reach , the activation of the first neuron in the lecture: the rules give . That slope is at and fades towards as grows, which is why a neuron driven deep into saturation barely learns — whatever gradient arrives at it is multiplied by almost nothing.
For our curve, the chain rule turns into , and the rest is term by term:
That formula is what the dashed tangents on this page are drawn from.
Where , the tangent is horizontal. On the plotted range that happens 7 times; where changes from negative to positive the curve has a local minimum — the deepest at , where — and where it changes from positive to negative, a local maximum.
Not every function has a derivative everywhere. has slope to the left of and to the right; the secant has no single limit at , so is not differentiable there. The activation function ReLU, , has the same corner, and neural-network libraries simply agree to use there.
Zoom into the corner and see for yourself. The panel draws two secants from the point, one to the left and one to the right, with tied to the zoom. On a smooth curve, zooming in straightens the picture and the two secants fold into one line. At the corner of the picture never changes: at every zoom it is the same V, the left secant stays at and the right at . There is no single number for the derivative to be.
The corner of |x| under magnification: the left and right secants from x = 0 keep slopes −1 and +1 at every zoom, while on the smooth x² they merge into one line; ReLU shows the 0 that libraries use at its corner.
How small should h be on a computer?
In exact arithmetic, a smaller is always better. On a computer it is not.
The secant slope , the forward difference, subtracts two nearly equal numbers. A 64-bit float carries about 16 significant digits, so the difference loses as many of them as and share. The rounding error that is left is then divided by the tiny .
Two errors pull in opposite directions:
- truncation error — the curve is not a line — is about and shrinks with ;
- rounding error is about , where is the float precision, and grows as shrinks.
Their sum is smallest near .
The central difference uses points on both sides. The terms cancel, its truncation error is of order , and its best moves up to about .
The panel computes both formulas in ordinary 64-bit arithmetic and plots their error against ; drag across it to move .
Log-log plot of the error of the forward and central differences of f at x = 3 against h: both errors fall as h shrinks until rounding takes over and they rise again, the forward difference bottoming out near h = 1e-8 and the central one near 1e-5.
Read the plot from right to left, the way shrinks. At first both lines fall with slopes 1 and 2: every tenfold decrease of divides the forward error by 10 and the central error by 100. Then rounding takes over and both lines turn up into noise.
At the forward difference is most accurate at , with error . The central difference reaches at — about 450 times more accurate, with 1000 times larger.
At both estimates are exactly : the float nearest to is itself, so .This is why numerical derivatives are used to check analytic ones — a “gradient check” with the central difference and around — and not to train networks: each estimate costs extra function evaluations per input, and the answer is good to only about ten digits.
The chain rule on a computation graph
Real formulas are compositions of simple operations. Break
into steps: , , . This is a computation graph.
The forward pass computes values left to right. With : .
The backward pass answers a different question: for every node, how much does change when that node is nudged? That is the derivative , called its gradient. A node’s value and its gradient are different numbers.
Each operation only needs to know its own local derivative:
- for : and ;
- for : — addition passes the gradient through unchanged;
- for : and — multiplication hands each input the value of the other.
Rates along a chain multiply. The lecture borrows George F. Simmons’s example: if a car travels twice as fast as a bicycle, and the bicycle four times as fast as a walking man, the car travels times as fast as the man. The chain rule multiplies local derivatives along the path in exactly that way:
Move the backward-pass slider: every notch processes one operation, from right to left, and labels each edge it crosses with its local derivative. Then drag any input sideways.
The computation graph of L = (a·b + c)·s with each node's value and gradient; a backward-pass slider processes one operation at a time from the output towards the inputs, labelling edges with local derivatives, and dragging an input compares the change in L with the gradient's prediction.
It starts from : nudging the output by changes it by . At the end , and dragging checks them the old way. Drag from to and moves by — exactly . The gradient is that prediction: how much the output moves per unit nudge of the input. Here it is exact, because is linear in each input on its own; for a curved function it holds only for small nudges, which is the whole content of .
Now move all four inputs at once. “Step along the gradient” adds to every input times its own gradient, and rises from to . The gradient predicted a rise of ; the true one is . The gap is the products of inputs — , — that change together and that a first-order prediction cannot see: is linear in each input alone, not in all of them at once. Stepping against the gradient would lower instead, and that is the whole of training, as the last section shows.
A ”—” before a node is reached means “not computed yet”, which is not the same as zero. Drag to and every gradient to its left becomes zero, although none of those nodes has value zero: when the last factor is zero, nothing upstream can move .
This is backpropagation. One backward sweep yields the gradient of the single output with respect to every node, at a cost of a small constant times the forward pass. A network has one loss and millions of parameters, and that ratio is why it trains backwards and not forwards.
One variable, several paths: gradients add
What if a value is used more than once? Take
The variable enters the graph three times: twice as an input of and once in . Each use is a separate path to , and each path sends back its own contribution: from each side of the product, from the second term. The multivariable chain rule says the derivative is their sum:
That is why a backpropagation engine writes grad += ... and not grad = .... The second button replays the bug: each contribution overwrites the one before, only the last survives, and the tangent drawn from that gradient misses the curve.
The graph of L = x·x + k·x, where x reaches L along three paths and sends back three contributions; next to it the curve L(x) with a tangent drawn from the accumulated gradient, which misses the curve when contributions overwrite instead of add.
Drag the point to with . The sum is zero and the tangent is horizontal, yet no path has disappeared: . Contributions can cancel, and a zero gradient does not mean nothing depends on .
Two different things are called “accumulating” gradients, and it pays not to mix them up. Within one backward pass, contributions must add. Between training steps, the stored gradients must be reset to zero first — otherwise the previous step’s gradient is added to the new one. The first is the chain rule; forgetting the second is a bug.
Order matters too. A node may pass its gradient on only after every use of it has reported back, which is why the backward pass visits nodes in reverse topological order.
What the derivative is for: gradient descent
The sign of says which way goes up. To go down, step the other way:
The step size is the learning rate. Where the slope is steep, the step is long; near a minimum, where , it is short, and the walk stops by itself.
Gradient descent on the wavy curve: a point repeatedly steps against the slope of its tangent until the slope is zero at a local minimum; a large learning rate makes it bounce.
From with , the walk reaches the deepest valley, , in 20 steps. Drag the start (or move its slider) to instead — just past the peak at — and the dashed preview shows the same rule taking it the other way, into the shallower valley at . The derivative is local information: it knows which way is down here, not where the best valley is.
Now raise . Near the minimum at the curvature is , and the walk settles only if . Above that, every step overshoots by more than it corrects, and the point bounces between the walls of the valley instead of settling.
Training a neural network is this same loop. becomes millions of weights, becomes the loss on the training data, becomes the gradient vector, and backpropagation computes it in one sweep. The rest of the lecture builds exactly that on top of the tiny engine in the panels above.
Short answers
What is a derivative in one sentence?
The derivative f′(x) is the rate at which f changes per unit change of its input near x — the slope of the straight line that best approximates the graph at that point.
Why is the tangent defined by a limit?
A single point does not determine a line's slope: infinitely many lines pass through it. A second point at distance h does determine one, the secant, and the tangent is what the secant turns into as h goes to zero.
Does every function have a derivative?
No. |x| has a corner at 0: the slope is −1 on the left and +1 on the right, so the secant has no single limit. ReLU, max(0, x), has the same corner; neural-network libraries simply pick a value there, usually 0.
What is the difference between a derivative and a gradient?
A derivative is the rate of change of a function of one variable. A gradient collects the partial derivatives of a function of many variables, one per input, into a vector; it points in the direction of steepest increase.
Numerical, symbolic or automatic differentiation — which is which?
Numerical differentiation estimates the slope from function values at nearby points and trades truncation error for rounding error. Symbolic differentiation manipulates formulas. Automatic differentiation, which backpropagation is an instance of, applies the chain rule to the actual sequence of operations a program executed and gives exact derivatives up to floating-point rounding.
Why does backpropagation run backwards?
A network has one loss and millions of parameters. Going backwards from the loss, one sweep over the graph yields the derivative of that single output with respect to every input, at a cost of a small constant times the forward pass. Going forwards would take one sweep per parameter.
Why are gradients accumulated with += in backpropagation?
Because a value that is used in several places influences the output along several paths, and by the multivariable chain rule its derivative is the sum of the contributions of all those paths. Assigning instead of adding keeps only the last path.
Further reading
- A. Karpathy, “The spelled-out intro to neural networks and backpropagation: building micrograd”, Neural Networks: Zero to Hero, lecture 1, 2022. The lecture this page follows.
- 3Blue1Brown, Essence of Calculus, video series, 2017. The geometric intuition behind derivatives and the chain rule.
- M. Spivak, Calculus, 4th ed., Publish or Perish, 2008. Chapters 9–10: the definition of the derivative and the rules, proven properly.
- J. Nocedal, S. J. Wright, Numerical Optimization, 2nd ed., Springer, 2006. Chapter 8: finite differences, their errors, and automatic differentiation.
- A. G. Baydin, B. A. Pearlmutter, A. A. Radul, J. M. Siskind, “Automatic differentiation in machine learning: a survey”, Journal of Machine Learning Research 18 (2018), 1–43.
- A. Griewank, A. Walther, Evaluating Derivatives: Principles and Techniques of Algorithmic Differentiation, 2nd ed., SIAM, 2008. The reference on reverse-mode differentiation.