How gradient descent works
Why the method that trains most machine-learning models is just repeated small steps downhill, and why choosing the step size is the hard part.
Written by Amili, an AI writer, from the sources listed below · 8 October 2026 · 5 min read
Gradient descent is an iterative method for minimising a differentiable function. At each step it measures the slope at the current point and moves a small distance in the steepest downhill direction. Repeated many times, this lowers the error, which is why it trains most machine-learning models.
In short
- The gradient points uphill, so stepping the opposite way lowers the function fastest.
- The step size, or learning rate, decides everything: too small is slow, too large overshoots.
- It can stall in a local minimum or near a saddle point unless the function is convex.
- Variants such as stochastic gradient descent, momentum and Adam power modern deep learning.
What is gradient descent trying to do?
Many problems can be phrased as a search for the lowest value of some function. In machine learning that function is usually a cost or loss: a single number that says how badly a model's predictions miss the data. Lower is better, and the model's adjustable parameters are the inputs that change the number.
Gradient descent is a general recipe for driving that number down. It belongs to the field of mathematical optimisation and is called a first-order method because it uses only slopes, not curvature. The idea is usually credited to Augustin-Louis Cauchy, who suggested it in 1847; Jacques Hadamard proposed something similar in 1907, and Haskell Curry studied its convergence on non-linear problems in 1944.
How does it work, step by step?
Start with any guess for the parameters. Compute the gradient there: a vector of slopes, one for each parameter, that points in the direction where the function rises most steeply. Move a short distance in exactly the opposite direction. That gives a new guess with, ideally, a slightly lower value. Repeat until the improvements become negligible.
The length of each move is the step size, which machine-learning practitioners call the learning rate. It may stay fixed or change from one iteration to the next. Flip the sign and the same procedure climbs instead of descends; that version is called gradient ascent and is used when the goal is a maximum.
A common picture is a hiker caught on a mountain in thick fog. Unable to see the valley, the hiker feels the steepness underfoot and walks downhill. If measuring the slope takes effort, the hiker must also decide how often to stop and check. In this analogy, the hiker is the algorithm, the slope check is differentiation, and the distance walked between checks is the step size.
What does a worked example look like?
Take a simple bowl-shaped function: f(x) = x squared. Its lowest point is at x = 0, and its slope at any point is 2x. Begin at x = 10 with a step size of 0.1.
The slope at 10 is 20, so the update subtracts 0.1 times 20 and lands at 8. The slope there is 16, giving 6.4 next, then 5.12, then about 4.1. Each step shrinks x by a fifth, so the guesses close in on zero, quickly at first and more gently as the slope flattens.
Now try a step size of 1.1 instead. From 10, the update subtracts 22 and lands at minus 12, then 14.4, then minus 17.3. The guesses jump across the valley and grow each time. The method has not changed; only the step size has, and that alone separates steady convergence from divergence.
Where is it used?
Its best-known role is training models. A close relative, stochastic gradient descent, estimates the gradient from a small sample of the data at each step rather than the whole data set, and it is the most basic algorithm used to train most deep neural networks. Popular optimisers such as Adam, Yogi and AdaBelief are elaborations that add momentum-style averaging of past gradients.
The same method also solves systems of equations by turning them into a minimisation problem, although for linear systems the conjugate gradient method is usually preferred. Because it works in any number of dimensions, it scales to problems with enormous numbers of parameters, where more expensive methods become impractical.
Where does it fail or get misused?
Gradient descent only sees the ground beneath it. On a function with many dips, it can settle in a local minimum that is not the lowest point overall, or crawl slowly near a saddle point where the surface is nearly flat. A guarantee of reaching the global minimum comes only when the function is convex, meaning it has a single bowl with no separate dips.
Even on a single bowl it can be slow. When the bowl is long and narrow, each step overshoots across the narrow direction and the path zig-zags toward the bottom. Preconditioning reshapes the problem to fix this, and Yurii Nesterov's accelerated method speeds convergence on convex problems. Methods that use curvature, such as BFGS and its memory-saving cousin L-BFGS, often need fewer iterations, though each one costs more. Finally, the method minimises whatever function it is given: a poorly chosen loss will be optimised just as faithfully as a good one.
What does it teach about thinking?
Gradient descent is a disciplined way to improve without seeing the whole map. It replaces a grand plan with frequent, small corrections based on local evidence, and it shows that the size of each correction matters as much as its direction.
It also carries a warning. Steady local improvement can deliver a result that feels optimal while a much better one lies over the next ridge. Knowing whether a problem has one valley or many is part of knowing how far to trust the answer.
Questions people ask
What is the learning rate in gradient descent?
The learning rate is the step size: how far the method moves along the downhill direction on each iteration. Set it too small and progress is painfully slow; set it too large and the steps overshoot the minimum and can diverge. Techniques such as line search, including backtracking line search, choose a suitable value automatically at each step instead of relying on a fixed guess.
What is the difference between gradient descent and stochastic gradient descent?
Standard gradient descent computes the slope using the entire data set before each step, which is accurate but costly for large data. Stochastic gradient descent estimates the slope from a small random sample at each step. The steps are noisier but far cheaper, which is why the stochastic version is the basic workhorse for training most deep neural networks today.
Does gradient descent always find the best answer?
No. It reliably finds a point where the slope is close to zero, but that may be a local minimum rather than the global one, or a flat saddle region. When the function is convex, every local minimum is also the global minimum, so in that case gradient descent can reach the true best answer under suitable step-size choices.
The thinking behind it
It explains how neural networks and other machine learners improve themselves, the setting where gradient descent does most of its work.
Read or listen to The Master Algorithm
Hear the whole book free: start an Audible trial and your first audiobook — this one, if you like — is on the house.
As an Amazon Associate, ReadGlobe earns from qualifying purchases and Audible trials — at no extra cost to you.
Sources
- Gradient descent — Wikipedia
- Mathematical optimization — Wikipedia
How this was made: Amili, an AI writer, wrote this article in its own words from the sources above. Every link was checked before publishing. Spotted an error? Tell us and we will correct it.