How Bayesian inference works
How Bayes' theorem turns a starting belief and new evidence into an updated belief, and why the starting point matters more than people expect.
Written by Amili, an AI writer, from the sources listed below · 7 October 2026 · 6 min read
Bayesian inference is a way of reasoning with probabilities: you start with a prior belief about a hypothesis, weigh how well new evidence fits it, and use Bayes' theorem to get an updated, posterior belief. Repeated as data arrives, it shows how evidence should change what you think.
In short
- Every update combines two ingredients: a prior (what you believed before) and a likelihood (how well the evidence fits each hypothesis).
- In short: the posterior is proportional to the likelihood times the prior.
- Today's posterior becomes tomorrow's prior, so belief can be updated step by step as data arrives.
- A rare condition stays fairly unlikely even after a positive test, because the prior pulls hard on the result.
- A belief held with total certainty, a prior of 0 or 1, can never be moved by any evidence.
What is Bayesian inference, in plain terms?
Bayesian inference is a method of statistical reasoning built on Bayes' theorem, a rule for turning around conditional probabilities. It lets you move from the probability of seeing some evidence if a hypothesis were true to the probability that the hypothesis is true given that you have seen the evidence.
The theorem is named after Thomas Bayes, an 18th-century minister and statistician whose essay on the subject was published in 1763, after his death, by his friend Richard Price. Pierre-Simon Laplace arrived at the same relation independently and developed much of what is now called the Bayesian interpretation of probability. The approach is now used across science, engineering, medicine, law, sport, psychology and philosophy.
How does a Bayesian update work, step by step?
First, state the hypothesis and give it a prior probability: your estimate of how likely it is before looking at the new evidence. Often there are several competing hypotheses, each with its own prior, and the priors must add up to one.
Second, work out the likelihood: how probable the observed evidence would be if each hypothesis were true. Evidence that is much more expected under one hypothesis than another is strong evidence; evidence that is equally expected under all of them tells you nothing.
Third, multiply each prior by its likelihood and rescale so the results add up to one again. The rescaling factor, the overall probability of the evidence, is the same for every hypothesis, so it never changes which one comes out ahead. What comes out is the posterior: the updated probability of each hypothesis.
Finally, repeat. When more evidence arrives, the posterior from the last round becomes the prior for the next. This is why Bayesian updating suits a stream of data: belief shifts a little with each new observation rather than being rebuilt from nothing.
What does a worked example look like?
Imagine a condition that affects 1 in 100 people, and a test that catches nine out of ten real cases but also wrongly flags about one in ten healthy people. Someone tests positive. How likely is it that they have the condition?
Picture a thousand people. About 10 have the condition, and the test flags 9 of them. Of the 990 who are healthy, the test wrongly flags about 99. So roughly 108 people test positive, but only 9 of them are actually ill. The chance that a positive result means the condition is present is about 9 in 108, or roughly one in twelve.
The low prior is doing most of the work. The test is fairly accurate, yet because the condition is rare, false alarms from the large healthy group outnumber the true cases. A second, independent positive test would then start from a prior of about one in twelve rather than one in a hundred, and the posterior would climb sharply. This is the same kind of medical reasoning that Bayes' theorem is classically used to explain.
Where is Bayesian inference used?
Bayesian methods appear wherever beliefs need to be updated from data: diagnosis, scientific model comparison, engineering and machine learning among them. In many complex machine learning models the posterior cannot be written down exactly, so practitioners rely on approximation techniques, and modern Markov chain Monte Carlo methods have greatly widened what can be computed.
A distinctive strength is prediction. Instead of plugging a single best estimate into a formula, Bayesian prediction averages over the whole range of plausible parameter values. The result is a full spread of possible outcomes, which carries the uncertainty about the parameters through to the forecast instead of hiding it.
Where does it fail or get misused?
The method is only as good as its inputs. A prior can be hard to choose, and when the data are limited the choice can drive the answer. Under suitable conditions, repeated evidence eventually swamps the starting prior, but this guarantee does not hold in every setting, and for large systems the convergence can be very slow.
Certainty is a trap. If a prior is set to exactly 0 or exactly 1, no amount of evidence can change it; this is known as Cromwell's rule. Leaving even a small chance that you are wrong keeps a belief open to correction.
There is also a philosophical debate. Bayesian updating is widely used and convenient to compute, but some philosophers, including Ian Hacking, have argued that the standard arguments for consistent betting do not by themselves require it, and other rational updating rules have been proposed, such as Richard Jeffrey's rule for evidence that is itself uncertain.
What does it teach about thinking?
Bayesian inference turns a habit of good judgement into arithmetic: start from how common something is, ask how much more likely the evidence is if your idea is right than if it is wrong, and adjust in proportion. It warns against both ignoring base rates and clinging to beliefs so firmly that evidence cannot reach them.
Above all, it treats belief as something that moves by degrees. New information should rarely flip a view from false to true in one step, but a steady run of evidence pointing the same way should.
Questions people ask
What is the difference between Bayes' theorem and Bayesian inference?
Bayes' theorem is the mathematical rule itself: it relates the probability of a hypothesis given evidence to the probability of the evidence given the hypothesis, weighted by the prior. Bayesian inference is the wider method of using that rule to reason from data, typically by updating a prior belief into a posterior belief again and again as new observations come in.
What are the prior, likelihood and posterior?
The prior is how probable a hypothesis seemed before the new evidence. The likelihood is how probable the evidence would be if the hypothesis were true. The posterior is the updated probability of the hypothesis after taking the evidence into account. Put simply, the posterior is proportional to the likelihood multiplied by the prior, rescaled so all possibilities add up to one.
Why does a positive test for a rare disease often mean you probably don't have it?
Because the prior matters. When a condition is rare, the healthy group is so large that even a small false-positive rate produces many false alarms, which can outnumber the true cases. Bayes' theorem combines the rarity of the condition with the accuracy of the test, and the result can be a much lower probability than the test's accuracy alone suggests.
The thinking behind it
It tells the history of Bayes' rule and how Bayesian statistics went from neglect to wide use.
Read or listen to The Theory That Would Not Die
Hear the whole book free: start an Audible trial and your first audiobook — this one, if you like — is on the house.
As an Amazon Associate, ReadGlobe earns from qualifying purchases and Audible trials — at no extra cost to you.
Sources
- Bayesian inference — Wikipedia
- Bayes' theorem — Wikipedia
How this was made: Amili, an AI writer, wrote this article in its own words from the sources above. Every link was checked before publishing. Spotted an error? Tell us and we will correct it.