The multi-armed bandit: how algorithms decide when to explore and when to exploit

Keep ordering your favourite dish, or try something new? The multi-armed bandit turns that everyday dilemma into mathematics, and the answers power trials, routing and recommendations.

Written by Amili, an AI writer, from the sources listed below · 6 October 2026 · 5 min read


The multi-armed bandit is a decision problem: you repeatedly choose between options whose payoffs you only partly know, and you want the biggest total reward. Every choice trades exploiting the option that looks best against exploring others that might be better. Algorithms such as epsilon-greedy, Thompson sampling and upper confidence bounds manage that trade.

In short

  • The name comes from a gambler facing a row of slot machines, once nicknamed one-armed bandits.
  • Exploiting uses what you know; exploring buys information you might use later.
  • Success is measured by regret: how much reward you gave up compared with always picking the best option.
  • Simple strategies such as epsilon-greedy mostly exploit but deliberately explore at random now and then.
  • The model is used for clinical trials, network routing, portfolio design and new-item recommendations.

Where does the name come from?


Imagine a gambler standing in front of several slot machines. Each machine pays out according to its own hidden odds. The gambler has a limited number of pulls and wants to walk away with as much money as possible. Which machines should they play, how often, and when should they stop playing one and try another?

That picture gave the problem its name, since slot machines were once called one-armed bandits. In the general version, a decision maker picks one of several fixed options again and again, learns a little about each option every time it is chosen, and tries to collect the largest total reward. An important detail is that choosing an option does not change that option: the machine's odds stay the same no matter how often it is played.

What is the explore–exploit trade-off?


Every pull is a choice between two goals. Exploitation means playing the machine that has paid best so far, cashing in on current knowledge. Exploration means trying a machine you know less about, accepting a likely lower reward now in exchange for information that could pay off later.

Lean too hard on exploitation and you may lock onto a mediocre machine early, never discovering that the one next to it was better. Lean too hard on exploration and you waste pulls on machines you already know are poor. The art lies in the balance, and the right balance depends on how many pulls remain: with a long horizon ahead, information is worth more, so exploring makes more sense than it does near the end.

How do bandit algorithms work?


Researchers score a strategy by its regret: the gap between the reward an all-knowing player would have collected by always choosing the best option and the reward the strategy actually collected. A good strategy keeps regret low, and a strategy whose average regret per round shrinks towards zero over time will, given enough rounds, settle on the best choice.

The simplest popular method is epsilon-greedy. Most of the time it picks whichever option has the best average so far. With a small probability, called epsilon, it picks an option at random instead. Upper confidence bound methods are more deliberate: they favour options whose payoff is either high or still uncertain, so that poorly tested options get a fair hearing. Thompson sampling chooses options in proportion to how likely each one is to be the best, given the evidence so far.

For one classic version, with rewards discounted over time, John Gittins proved that a single number per option, now called the Gittins index, is enough to find the optimal policy.

A worked example: choosing a lunch spot


Say there are five places to eat near your office and you will have lunch out a hundred more times this year. After trying two of them once each, one was decent and one was disappointing. An exploit-only rule says: go to the decent one forever. An epsilon-greedy rule says: go to the decent one most days, but on roughly one day in ten, pick at random, which eventually gets the three untried places a visit.

If one of those untried places turns out to be excellent, the few disappointing experiments along the way were a small price for many better lunches to come. If you only had two lunches left, the same experiments would rarely be worth it. That is the horizon effect in miniature.

Where is it used, and where does it struggle?


The bandit model fits any setting where you must keep acting while still learning. It has been applied to clinical trials that compare treatments while trying to limit harm to patients, to adaptive routing that reduces network delays, to financial portfolio design, and to deciding which research projects a large organisation should fund. Recommender systems use bandit methods to handle the cold-start problem, when a new item has too little data to judge. A variant called best arm identification, which only asks which option is best by the end, is used in A/B testing.

The method struggles when its assumptions break. Rewards that come rarely give a learner little to go on. Rewards that look good early but lead nowhere can lure it away from better long-term options. And in a full reinforcement-learning setting, where actions change the world, the simple bandit picture no longer holds. The problem was hard enough that, according to the statistician Peter Whittle, Allied scientists in World War II joked about dropping it over Germany so that enemy scientists could waste their time on it too. Herbert Robbins formulated the version most studied today in 1952.

What it teaches about everyday decisions


The bandit problem gives a precise reason why curiosity is rational. Trying something new is not a lapse in discipline; it is an investment in information, and it is worth most when you have a long time ahead to use what you learn.

It also explains why settled habits make sense later in a process. As the horizon shortens, exploiting what you know becomes the better bet. A useful question before any repeated choice, from restaurants to suppliers to hiring channels, is how many more times the decision will come up, and what finding out is worth.

Questions people ask


What is regret in a multi-armed bandit?

Regret is the expected difference between the total reward of the best possible strategy, which always picks the option with the highest average payoff, and the total reward a strategy actually collected. Lower regret means less was lost while learning. Strategies whose average regret per round tends to zero eventually behave like the best strategy.

What is the difference between epsilon-greedy and Thompson sampling?

Epsilon-greedy usually takes the option with the best average so far and occasionally, with probability epsilon, picks one at random. Thompson sampling instead chooses each option with the probability that it is the best one, given the evidence collected, so its exploration is guided by uncertainty rather than by a fixed random rate.

How is a multi-armed bandit different from an A/B test?

A classic A/B test splits traffic in fixed proportions until the end and then picks a winner. A bandit keeps shifting choices towards the better option while the test is still running, which reduces the reward lost along the way. Best arm identification, a bandit variant focused only on naming the winner, is closest to A/B testing.

The thinking behind it


Brian Christian and Tom Griffiths devote a chapter to explore and exploit, applying bandit thinking to everyday choices such as where to eat.

Read or listen to Algorithms to Live By

Hear the whole book free: start an Audible trial and your first audiobook — this one, if you like — is on the house.

As an Amazon Associate, ReadGlobe earns from qualifying purchases and Audible trials — at no extra cost to you.

Sources

How this was made: Amili, an AI writer, wrote this article in its own words from the sources above. Every link was checked before publishing. Spotted an error? Tell us and we will correct it.

More algorithms, explained


Books readers reach for

As an Amazon Associate, ReadGlobe earns from qualifying purchases — at no extra cost to you.