In short
Run an A/B test for at least two full weeks and until each variant has reached the sample size calculated before the test began, which for a typical conversion rate of 2–5% and a hoped-for lift of 20% means several thousand visitors and a few hundred conversions per variant. Do not stop early because one variant is ahead. A page receiving fewer than roughly 1,000 conversions a month cannot detect realistic improvements in a reasonable time and should improve through research, fixes and sequential measurement instead of testing.
The two rules
Rule one: decide the sample size before you start, and run until you reach it. An A/B test compares two conversion rates and asks whether the difference is larger than random variation would produce. How many visitors that takes depends on the baseline rate and the size of improvement you want to detect. Any free sample-size calculator will give the number; the point is to compute it first and treat it as the finish line.
Rule two: run for whole weeks, at least two. Visitors behave differently on Tuesdays and Saturdays, at the start of the month and the end. A test that runs Monday to Thursday has sampled a different population from the one that will see the winner. Two weeks is the minimum to cover the cycle twice; four is more typical once the sample size is factored in.
Stopping when one variant "looks like it is winning" violates both, and it is the most common way tests produce confident wrong answers.
What the sample size really is
To make the numbers concrete:
| Baseline conversion rate | Lift you want to detect | Visitors needed per variant (approx.) |
|---|---|---|
| 2% | 10% relative (2% to 2.2%) | ~78,000 |
| 2% | 20% relative (2% to 2.4%) | ~20,000 |
| 5% | 20% relative (5% to 6%) | ~7,500 |
| 10% | 20% relative (10% to 12%) | ~3,500 |
| 10% | 50% relative (10% to 15%) | ~600 |
Figures at 95% confidence and 80% power, rounded; calculators vary slightly. Two things stand out. Small lifts on low conversion rates need enormous traffic, which is why most published wins come from high-traffic sites. And a page converting at 2% needs around 40,000 visitors to detect a solid 20% improvement, which for many business sites is a year of traffic to that page.
Why stopping early is worse than not testing
Conversion rates wander. In the first few days of a test, one variant will be ahead simply by chance, often by a wide margin, and a dashboard will happily colour it green. If you stop there, you ship whichever variant got lucky first. Simulations of this behaviour show that stopping at the first sign of significance produces a false winner far more often than the stated confidence level suggests, sometimes most of the time.
The damage is not just a wasted test. A false winner becomes an organisational belief: "we tested it, the green button won", and the belief outlives the evidence.
Is your site too small?
A rough threshold: if the page you want to test receives fewer than about 1,000 conversions a month, most realistic tests will take too long to be worth running. Below a few hundred conversions a month, testing is essentially unavailable, because the only lifts you could detect are so large that you did not need a test to see them.
Signs you are below the line: your test calculator says months; your previous tests all ended "inconclusive" or were stopped early; the conversion you care about happens a few times a day. This describes most service businesses, most B2B sites and most local sites, and it is not a failing. It means a different method.
What to do instead with low traffic
Fix the obvious first. Slow pages, unclear headlines, forms with too many fields, a missing call to action. These are known improvements that do not need testing; they need doing.
Research, then change, then measure sequentially. Talk to five customers. Watch session recordings. Read the search queries bringing people to the page. Make the change the evidence points to, then compare the following month's segmented conversion rate with the previous one. It is not a controlled experiment, but with a clear before-and-after and no other changes in the period, it is a defensible inference and it is what small sites have.
Test bigger changes. If you must test, test something that could plausibly move conversion by 50%, such as a completely different page structure or offer, rather than a button colour. Large effects need small samples.
Test higher up the funnel. A page with 200 conversions a month may have 20,000 visitors. A test whose success metric is clicks to the next step, rather than final conversions, has a hundred times the sample. Just be honest that a click is a proxy.
Pool across pages. A template change tested across all product pages accumulates conversions faster than one page alone.
If you do have the traffic
Write a hypothesis with a predicted direction and size. Compute the sample. Run for whole weeks and do not peek in a way that changes the decision. Check that the variants received similar traffic and that no campaign or outage skewed the period. Report the result with its confidence interval, including when it is a loss or a null, and archive it so the next person does not re-run it. Then move on: a well-run test that changes nothing is still knowledge.
If you are not sure which side of the threshold a site is on, or want help setting up sequential measurement that you can trust, book a call.
Common questions
What is the minimum time to run an A/B test?
Two full weeks, to cover weekday and weekend behaviour twice, and longer if the pre-calculated sample size has not been reached. Most tests on mid-traffic sites run three to six weeks. Ending sooner because one variant is ahead produces false winners more often than the stated confidence suggests.
How many visitors do you need for an A/B test?
It depends on the baseline conversion rate and the lift you want to detect. At a 2% conversion rate, detecting a 20% relative improvement needs roughly 20,000 visitors per variant; detecting 10% needs nearly 80,000. At a 10% rate, a 20% lift needs about 3,500 per variant. Use a sample-size calculator before starting and treat the result as the finish line.
Can you A/B test with low traffic?
Rarely with confidence. Below roughly 1,000 conversions a month on the tested page, realistic tests take months; below a few hundred, only very large effects are detectable. Low-traffic sites do better with customer research, session recordings, fixing known problems, and measuring conversion month over month after each change.
What is statistical significance in A/B testing?
It is the probability that a difference as large as the one observed would appear by chance if the variants were really identical, conventionally set at 5%. It does not mean the effect is large or important, and it is only valid if the test ran to its planned sample without early stopping. Report the confidence interval alongside it.
Why do most A/B tests fail?
Most well-run tests end in no significant difference, because most changes do not matter much; that is a legitimate result. Tests that produce misleading winners usually failed procedurally: stopped early, run over an incomplete weekly cycle, skewed by a campaign, or testing a change too small for the available traffic.
