azyware
Personalisation & machine learningTechnique / practice

A/B test (controlled test)

Also: split test, randomised controlled experiment

In one sentence

What is A/B test (controlled test)?

An A/B test randomly assigns users to a control group and one or more variants, then compares outcomes, so a change's effect on revenue, conversion or retention is measured rather than assumed.

What A/B test (controlled test) means

An A/B test is the only reliable way to know whether a personalisation change worked. Users are randomly split; the control group sees the existing experience and the variant group sees the new ranking, model or message. Because assignment is random, the groups differ only in the treatment, and any difference in outcome can be attributed to it. The test runs until enough users have been observed for the result to be statistically distinguishable from noise, and the primary metric is fixed before it starts.

For personalisation the practical details matter: assignment must be stable per user across sessions and channels, the metric should be a business outcome such as revenue per visitor rather than click-through, and the test should run through at least one full weekly cycle. Guardrail metrics such as returns, latency and support contacts catch side effects.

It is not a before-and-after comparison, which is confounded by seasonality, campaigns and everything else that changed. It is also not a demo that looks better. The result of a controlled test is the uplift figure a personalisation programme is judged on.

Who it really matters to

  • Founder / CEO: it is how you know whether a personalisation vendor delivered anything; insist on the control group before the launch, not after.
  • CFO: a measured uplift with a confidence interval is a number you can put in a business case; a screenshot of a nicer homepage is not.
  • Product manager: fixing the primary metric and sample size in advance prevents the temptation to stop early or pick the metric that moved.
  • Data lead: randomisation, exposure logging and analysis are your domain; a test with leaky assignment or peeking is worse than no test.

Why it exists

Controlled tests exist because personalisation results are easy to fake, even unintentionally. Launch a new recommender in November and revenue rises; was it the model or the festive season? Compare users who clicked recommendations with those who did not and the clickers buy more; but they were already the engaged ones. Randomisation removes those confounders. The trade-off is patience and traffic: a test needs enough users and enough time, and during it a share of visitors gets the old experience. That cost is small next to scaling a change that never worked.

Where it is applied

  • A D2C brand testing a new product-page recommender against the existing bestseller carousel on revenue per session.
  • A SaaS company testing personalised onboarding checklists against the generic flow on trial-to-paid conversion.
  • A bank testing personalised versus generic push notifications on activation of a new card, with complaint rate as a guardrail.
  • A retail chain testing a demand-forecast-driven replenishment rule in a random subset of stores on stock-outs and waste.
  • An ed-tech platform testing adaptive lesson ordering against the fixed syllabus on completion rate.

Is A/B test (controlled test) a skill?

Technique / practiceAn experimental technique and the measurement standard for personalisation work. Every Eazyware personalisation engines engagement is accepted on a controlled test result, in line with the evals-over-demos stance.

Eazyware service that covers it: Personalization Engines. Starting prices are on the pricing page.

Frequently asked questions

How long does an A/B test for personalisation need to run?

Long enough to reach the sample size fixed in advance and to cover at least one full weekly cycle, commonly two to four weeks for a mid-size retailer. Stopping when the result first looks good produces false wins.

What if we cannot afford to show the old experience to half our users?

The control group can be smaller, such as ten percent, at the cost of a longer test. What you cannot afford is scaling a change whose effect you never measured.

Related reading

Need A/B test (controlled test) built, not just explained?

PRJECT IN MIND?