A/B Tests: Hypotheses, Sample Size, Metrics, and Stop Rules

An A/B test randomly assigns eligible units to a control or treatment experience and estimates what the assigned change did to a chosen outcome. The arithmetic is usually the easy part. The result becomes trustworthy only when the hypothesis can be disproved, assignment and telemetry hold up, the metrics answer a real decision, and the team decides in advance how much evidence is enough.

A/B tests: a large monitor showing an abstract split chart, balance scale, clock, target with dart, closed notebook, pen, coffee cup

The two letters are the least important part. Variant A is usually the current experience and variant B the proposed change, but a random split does not rescue an ambiguous treatment, a broken assignment, an opportunistic metric, or a test that stops when the dashboard looks favorable.

For a binary outcome, the point estimate is straightforward:

control rate = control conversions ÷ control assigned units
treatment rate = treatment conversions ÷ treatment assigned units
absolute effect = treatment rate − control rate
relative lift = (treatment rate − control rate) ÷ control rate

Here is illustrative arithmetic, not a company result. If 10,000 assigned control users produce 800 conversions, the control rate is 8.0%. If 10,000 assigned treatment users produce 860, the treatment rate is 8.6%. The point estimate is +0.6 percentage points, or 7.5% relative lift.

That calculation is not a ship decision. It still needs the planned uncertainty interval, sample-size target, metric and telemetry checks, guardrails, and stopping rule.

A useful hypothesis tells you what happens next

“Changing the page will increase conversion” is too loose. It does not name the eligible population, mechanism, treatment, primary outcome, time horizon, or action after the result.

A testable contract answers:

  • Population: Which units can enter, and when?
  • Assignment: What unit is randomized—person, account, device, session, geography, or time block?
  • Treatment: What exactly differs, and is that difference stable during exposure?
  • Mechanism: Why could the change move the outcome?
  • Primary metric: Which one result governs the decision?
  • Guardrails: Which harms can stop or block the change?
  • Minimum effect: What improvement is large enough to matter?
  • Decision rule: What will the team do for a positive, negative, inconclusive, or invalid result?

The mechanism matters because it predicts diagnostic movements. If a shorter form is expected to reduce friction, starts and completions may change in a particular sequence. If completion rises but qualified downstream outcomes fall, the mechanism did not produce the intended business result.

A hypothesis is not a prediction that treatment will win. It is a falsifiable explanation linked to a precommitted decision, including what the team will do when the evidence is inconclusive.

Randomize the unit that carries interference and memory

Random assignment supports causal interpretation when treatment and control groups are comparable except for the assigned change. The unit must match how exposure persists.

Session-level assignment can contaminate a test when the same person sees both variants and remembers the experience. User-level assignment can fail when several users in one account influence one another. Account-level assignment reduces that interference but also reduces the number of independent units. Geography or time-block assignment introduces different dependence and seasonality concerns.

Write the assignment key, allocation ratio, eligibility event, exposure event, exclusion rules, persistence window, and re-entry behavior before launch. Log assignment even when the experience fails to render, or treatment-dependent data loss can remove precisely the units most affected by the change.

Microsoft’s sample-ratio-mismatch research describes SRM as an unexpected difference between observed and configured assignment proportions. It can arise from assignment, triggering, redirects, telemetry loss, or filtering. The SRM test is a warning; investigation still has to find the cause.

Microsoft treats sample-ratio mismatch as a serious data-quality signal. A statistically favorable outcome is not trustworthy until the mismatch is explained and repaired or the test is declared invalid.

Metrics need different jobs

Microsoft’s during-experiment patterns separate overall outcome, feature and diagnostic, guardrail, and data-quality metrics. That taxonomy prevents one crowded dashboard from treating every movement as an equal reason to ship.

Metric jobQuestionExample form
Primary decisionDid the intended outcome improve enough?Qualified completion per assigned unit
GuardrailDid the change cause unacceptable harm?Error, latency, complaint, cancellation
DiagnosticDid the proposed mechanism occur?Start, step completion, feature use
Data qualityCan the other metrics be trusted?SRM, assignment loss, join rate, missingness

Use a denominator that treatment cannot silently redefine. “Purchases per checkout starter” can mislead if treatment changes who starts checkout. “Purchases per assigned eligible user” preserves the randomized population, while the starter rate can remain a diagnostic.

Declare how repeated events, bots, refunds, late conversions, missing data, outliers, currency, and account changes are handled. Archive the query or code version. A metric name is not a definition.

Sample size is planned from the decision, not copied from a rule of thumb

The A/B sample-size guide by Georgiev and colleagues explains that planning depends on baseline rate or variance, effect definition, significance, power, allocation, and the analysis. Correlated observations and absolute versus relative effects require additional care. NIST’s comparison reference likewise derives proportion-test requirements from the assumed rates and error probabilities.

Start with the minimum detectable effect, or MDE: the smallest change the experiment is designed to detect with the chosen error rates. The MDE should come from the decision. If a change smaller than 0.3 percentage points cannot repay implementation and operating cost, powering for 0.05 points wastes traffic. If a 0.3-point harm is unacceptable, the guardrail may need more precision than the primary upside metric.

No universal sample size or duration exists. Required independent units increase when the baseline is noisy, the worthwhile effect is smaller, higher power is required, the significance threshold is stricter, or allocation is uneven. Seasonality, novelty, learning, network effects, and conversion lag can require a longer calendar even after the numerical sample target is reached.

Plan both an information target and a calendar boundary. Reaching a user count in two hours does not observe a weekly retention outcome. Running for several weeks does not repair too few independent accounts.

Stopping rules determine whether repeated looks are valid

A fixed-horizon design selects its sample or duration and analysis before launch, monitors only for integrity and safety, and evaluates the primary result at the planned end. Repeatedly checking a conventional fixed-horizon p-value and stopping on the first favorable look changes its error behavior.

Sequential methods can allow continuous monitoring when the analysis is explicitly built for it. The always-valid inference paper develops one such statistical contract. Its existence does not make any real-time dashboard sequentially valid; the estimator, boundaries, and decision rule must implement the method actually claimed.

Safety stopping is different from success stopping. A severe reliability regression can end exposure early under a predeclared guardrail even if the primary analysis is fixed-horizon. Preserve the reason as “stopped for harm,” not “treatment lost under the primary metric.”

Write four end states before launch:

  1. Ship: primary benefit crosses the decision threshold and guardrails pass.
  2. Do not ship: evidence shows material harm or insufficient value under the rule.
  3. Inconclusive: the interval includes materially positive and negative outcomes.
  4. Invalid: assignment, exposure, telemetry, interference, or analysis failed.

Check whether the experiment deserves an interpretation

  1. Verify assignment and exposure — Check configured versus observed allocation, persistence, contamination, eligibility, and whether exposure logs exist independently of the outcome.
  2. Verify data quality — Inspect missingness, joins, event counts, bots, duplicates, late data, and metric observation units by variant.
  3. Apply the planned analysis — Use the predeclared population, metrics, estimator, multiplicity treatment, horizon, and stopping rule.
  4. Read effect and uncertainty — Report absolute and relative effects with intervals; compare them with the minimum worthwhile effect and harm boundary.
  5. Check mechanism and guardrails — Use diagnostics to explain the effect and guardrails to block a locally positive but globally harmful change.
  6. Record the decision and limits — Archive the hypothesis, configuration, code, exclusions, result, decision, and conditions under which the conclusion may not transfer.

External validity remains bounded. A result for eligible users during one period establishes an effect for that tested population and implementation under the design assumptions. It does not guarantee the same effect for new users, other countries, another season, a later product state, or a permanently scaled rollout.

Design the test around the decision it must support, then ask whether the data earned the right to answer it. A small, clean experiment may end inconclusively. A large experiment with broken assignment or telemetry has not become evidence by accumulating traffic.

Frequently asked questions

What is an A/A test, and when is it useful?

An A/A test randomly splits traffic between two identical experiences to check the experiment plumbing rather than choose a winner. Adobe Target recommends the pattern when a team is validating a new tool’s setup, implementation, or reporting; plan its sample and stopping rule in advance because identical arms can still produce a false positive. If a completed A/A test shows a persistent gap, pause product experiments and inspect assignment, exposure, telemetry, and analysis before treating the platform as calibrated.

How is an A/B test different from a multivariate test?

An A/B test compares one packaged experience with another, while a multivariate test crosses independently varied factors so their main effects and interactions can be estimated. In a full two-level factorial design, k binary factors require 2^k combinations before replication, so changing three binary page elements creates eight cells. Use that design when the interaction between elements matters and traffic can support every cell; otherwise, bundle the changes into one treatment and answer the simpler ship question.

Can multiple A/B tests run at the same time?

Concurrent tests can be valid when each assignment is independent and one treatment is unlikely to change how the other treatment works or is measured. Microsoft describes allowing users into multiple tests because isolating every experiment would divide traffic and reduce statistical power. Before launch, compare the tests’ population, surface, mechanism, primary metric, and guardrails; use mutual exclusion or an explicit factorial design when a plausible interaction would change the decision.

What does a p-value mean in an A/B test?

A p-value indicates how incompatible the observed data are with a specified statistical model; it is not the probability that the treatment works. The American Statistical Association also cautions that it does not measure effect size or practical importance. Read it alongside the estimated absolute and relative effects, uncertainty interval, minimum worthwhile effect, and integrity checks instead of turning a threshold crossing into the ship decision.

One person. A whole marketing team.

Invite only