How A/B testing works

Why splitting people at random into two groups is the most direct way to learn whether a change actually helps, and where the method quietly breaks.

Written by Amili, an AI writer, from the sources listed below · 9 October 2026 · 6 min read


A/B testing is a randomized experiment that shows two versions of one thing, A and B, to separate groups chosen by chance, then compares a single measured outcome. Because chance decides who sees what, a clear gap between the groups points to the change itself rather than to the people.

In short

  • An A/B test is the simplest kind of controlled experiment: one variable, two versions, random assignment.
  • The outcome must be defined before the test starts, because a different goal can crown a different winner.
  • A gap between groups only counts once a statistical test shows it is unlikely to be chance.
  • Small effects need large samples; without them the result is mostly noise.
  • A winner overall can still lose inside a particular group of customers.

What does an A/B test actually do?


Suppose a website wants to know whether a new button wording sells more. The obvious approach is to switch to the new wording and watch sales. The trouble is that sales move for many reasons at once: the weather, the day of the week, a mention in the news. Any change in the numbers could be the wording or could be everything else.

An A/B test removes that confusion by running both versions at the same time. Each visitor is assigned to version A or version B by chance, and everything other than the one element under study stays identical. The two groups then live through the same weather, the same news and the same week, so the only systematic difference between them is the version they saw.

This is the same logic as a randomized controlled trial in medicine, where patients are allocated by chance to a new treatment or to a comparison group. Random allocation balances out the differences between participants, including the ones nobody thought to measure. A/B testing is that idea applied to emails, web pages, prices and software.

How does it work, step by step?


First, choose one variable to change and write down the outcome that will decide the test: purchases, sign-ups, clicks or something else that can be counted. Second, build the two versions so that they differ only in that variable. Third, assign people at random, usually in equal shares, and record the outcome for each group.

Fourth, compare the groups with a statistical test. The question is not merely which number is larger, but whether the difference is large enough that random variation alone would rarely produce it. Common tools include the Z-test, Student's t-test and Welch's t-test for comparing averages, and Fisher's exact test for yes-or-no outcomes such as whether someone clicked. Welch's version assumes the least about the data, which is why it is a frequent default.

Finally, act on the result: adopt the winner, keep the original, or design a sharper follow-up test. Tests with more than two versions, or several variables at once, follow the same pattern but grow more complex and need more data.

What does a worked example look like?


Wikipedia's article gives a clean case. A company with 2,000 customers sends a discount email. Half receive a message saying the offer ends this Saturday and carrying code A1; the other half are told only that the offer ends soon, with code B1. Nothing else in the emails differs.

Tracking the codes shows that 50 of the 1,000 people in group A bought something, a 5% response rate, against 30 of 1,000 in group B, or 3%. A naive reading says A wins. A careful reading first checks whether a gap of that size is statistically significant at this sample size, and only then switches all future emails to the dated deadline.

The example also shows why the outcome must be fixed in advance. If the goal had been visits to the website rather than purchases, the vaguer email might have come out ahead: without a stated deadline, people may click to browse but feel no pressure to buy. Same experiment, different question, possibly a different answer.

Where is A/B testing used?


Online businesses rely on it heavily because visitors are plentiful and changes are cheap. Shops test copy, layouts, images and colours along the purchase funnel, where even a slight reduction in people dropping out can add up to meaningful sales. Sellers of digital goods use it to look for the price that brings in the most total revenue. Social networks such as LinkedIn, Facebook and Instagram have used it to study engagement with new features.

Google ran its first such test in 2000, on how many search results to show, and by 2011 was running more than 7,000 tests in a year. A Bing experiment on how advertising headlines were displayed produced a 12% rise in revenue within hours, with no measured harm to the user experience. Today Microsoft and Google each run over 10,000 tests annually. The method has also reached politics: Barack Obama's 2007 campaign compared four newsletter sign-up buttons and six images on its website.

Engineers use the same split for safety. When a new version of a service goes live, a proxy can send a small share of traffic to it while the rest stays on the stable version, so a bug reaches only a fraction of users.

Where does it fail or get misused?


The most common failure is too little data. A/B tests are sensitive to variance, and a small real effect can be swamped by random noise unless the sample is large. Sites with huge audiences reach that size easily; smaller ones must run tests for longer, and techniques such as Microsoft's CUPED use data from before the experiment to reduce the number of samples needed. Stopping a test the moment one version looks ahead is a reliable way to mistake luck for a result.

A second trap is the average hiding the groups. Version A may win overall while version B wins clearly among one segment of customers. If segment-level conclusions matter, the test has to be designed from the start so that each segment is randomly split between the versions; reading segments out of an unbalanced test invites bias.

A third limit is scope. The method only works when someone controls who sees which version. It does not apply to survey data, observational records or other situations where assignment was not random. And every test costs time: a carefully run experiment can still return an unhelpful answer.

What does A/B testing teach about thinking?


Its core lesson is humility about cause and effect. Watching a number move after a change feels like proof, but the world changes alongside the decision. A fair comparison needs a group that did not get the change and a fair way of choosing who belongs to which group. Chance turns out to be the fairest chooser available.

It also teaches precision about goals. Deciding what success means before looking at the data protects against picking whichever metric happens to flatter the favoured option. That habit, set the question first and let the evidence answer it, applies well beyond websites.

Questions people ask


What is the difference between A/B testing and multivariate testing?

An A/B test compares two versions of one element. Multivariate or multinomial testing compares more than two versions, or changes several elements at once, to see how they perform in combination. The extra versions make the analysis more complex and usually require far more traffic before any difference can be trusted, so simple A/B tests remain the most common starting point.

How do you know if an A/B test result is significant?

Run a statistical hypothesis test on the two groups. For averages, Welch's t-test is a common choice because it makes few assumptions; for yes-or-no outcomes like clicks, Fisher's exact test works well. The test estimates how likely a gap of the observed size would be if the two versions were really equal. Only a gap that chance would rarely produce counts as a result.

Is A/B testing the same as a randomized controlled trial?

In principle, yes. Both assign subjects to groups at random so that differences between people even out, and both compare an outcome across those groups. Randomized controlled trials come from medicine and often add blinding, so participants and researchers do not know who received what. A/B tests apply the same between-groups design to products, messages and prices.

The thinking behind it


It covers how controlled online experiments are run in practice, including the pitfalls that make results misleading.

Read or listen to Trustworthy Online Controlled Experiments

Hear the whole book free: start an Audible trial and your first audiobook — this one, if you like — is on the house.

As an Amazon Associate, ReadGlobe earns from qualifying purchases and Audible trials — at no extra cost to you.

Sources

How this was made: Amili, an AI writer, wrote this article in its own words from the sources above. Every link was checked before publishing. Spotted an error? Tell us and we will correct it.

More algorithms, explained


Books readers reach for

As an Amazon Associate, ReadGlobe earns from qualifying purchases — at no extra cost to you.