Three inputs decide it, and traffic is the output
An A/B test is a statistical procedure with a fixed appetite. Feed it less than it needs and it doesn't return a smaller answer — it returns an unreliable one, which is worse, because on screen an unreliable answer looks exactly like a reliable one.
Two of the four inputs are conventions you shouldn't be negotiating: 95% two-sided significance, meaning a 5% chance of calling a difference that isn't there, and 80% power, meaning a 20% chance of missing one that is. The other two are yours, and the second is where most requests fall apart.
- Your baseline conversion rate on the specific page or template you'd test — not the site average. A blog post converting at 0.3% and a pricing page at 6% are different experiments with wildly different appetites.
- The minimum detectable effect: the smallest improvement you'd actually act on, stated as a relative lift. "Any improvement" implies an infinite sample. Somebody has to name a number, and it should be the person who'd have to justify the change.
- How long you'll let it run. Weekly traffic times weeks is the entire budget. Two weeks minimum so you capture a full weekly cycle, and rarely beyond six before seasonality starts contaminating the comparison.
The number that matters is conversions, not sessions
Almost every discussion of this starts with traffic, which is why it stays confusing. Start with conversions instead and the arithmetic gets memorable.
- The conversions column barely moves with your conversion rate. Run it at 1% or at 6% and the figure stays within a few percent of what's above. A better-converting page doesn't need fewer conversions — it collects them from fewer sessions. That's the whole reason to hold conversions in your head instead of traffic.
- Sample scales with the square of the effect. Detecting +10% instead of +20% doesn't double the requirement, it roughly quadruples it. This is the most expensive decision in the exercise and it usually gets made in passing.
- Most conversion improvements that survive contact with reality land between 5% and 20% relative. Read the top three rows again with that in mind.
| Relative lift you want to detect | Conversions per variant | Sessions per variant at 2% | Total sessions for the test |
|---|---|---|---|
| +5% | about 6,200 | about 315,000 | about 630,000 |
| +10% | about 1,600 | about 81,000 | about 161,000 |
| +20% | about 415 | about 21,000 | about 42,000 |
| +30% | about 190 | about 9,800 | about 20,000 |
| +50% | about 75 | about 3,800 | about 7,600 |
Four ordinary setups, run through the same formula
The businesses below are illustrative, not clients. The arithmetic is real, and the pattern in it is the point of the page.
- Rows one and two: don't test. Not "test carefully" or "test for longer" — don't. A twelve-month experiment isn't an experiment, because the site, the market and the traffic mix all change underneath it.
- Row three is the surprise. It has the second-smallest traffic figure in the table and it's the only mid-size site that can test, because a 6% baseline collects the required conversions from far fewer sessions. High-intent B2B pages can start testing much earlier than blog traffic can.
- Row four is the argument. Same business, same page, same conversion rate, ten times the traffic — twelve months becomes five weeks. Nothing about the test changed. Only the input that SEO produces.
| Setup | Baseline | Monthly sessions on that template | Sessions needed in total | Time to finish |
|---|---|---|---|---|
| D2C skincare brand, product page | 1.2% | 6,000 | about 71,000 | about 12 months |
| Interior design studio, enquiry form | 2.5% | 2,000 | about 34,000 | about 17 months |
| B2B software, pricing page | 6% | 4,000 | about 13,000 | about 3 months |
| The same D2C brand, four years later | 1.2% | 60,000 | about 71,000 | about 5 weeks |
What calling it early actually does
Every testing platform shows a leading variant and a confidence figure from around day three. The arithmetic is correct and the conclusion is worthless, because the 5% error rate you signed up for assumes you looked once, at the end.
Check daily for a fortnight and you've taken fourteen chances to catch a 5% event. Your real false-positive rate is no longer 5%. How far above depends on how often you look and how quickly you stop, and it is not a small correction — it's the entire reason the sequential-testing literature exists, and why researchers built always-valid p-values for people who can't stop watching.
The failure mode has a recognisable shape. A test is stopped on day four with the variant 18% ahead. It ships. The metric drifts back to where it started over the following six weeks and nobody notices, because there's a new test running and the comparison window has closed. Worse, the phantom win hardens into a belief — "our audience responds to urgency" — and the next three tests get built on top of it.
- Write the required sample and the end date down before launch, in the same document as the hypothesis. If it isn't written before, it will be negotiated after.
- Run whole weeks. Stopping mid-week weights the sample towards whichever days happened to be included, and weekday and weekend buyers behave differently on almost every Indian site we've looked at.
- Name one primary metric in advance. Reading five and reporting the one that moved is peeking with extra steps.
- If you genuinely must monitor continuously, use a platform doing sequential or always-valid inference, and accept that it asks for more sample in exchange for letting you look.
What to do when the arithmetic says no
The wrong response is to run the test anyway and describe the outcome as directional. There's no such thing as a directional result from an underpowered test — there's a number, and it's roughly as likely to point the wrong way as the right one.
The right response is to stop using the word test and go and get information some other way. None of the following produces evidence in the statistical sense, and all of it produces better decisions than noise does.
- Five customer calls. People who bought last month, fifteen minutes each. Ask what nearly stopped them. Two of the five will say the same thing, and that's your next change.
- The objection log your sales team keeps in their heads. Write it down for a fortnight. The top three objections belong on the page, answered in the customer's words.
- Session replay on abandonment only. Watch twenty sessions that reached the form and left. You're hunting breakage, not inspiration.
- Internal site search queries. People typing into your search box are telling you exactly what the navigation failed to show them — how to read that data is a page of its own.
- One exit question. "What stopped you today?", free text. Fifty answers is plenty and it takes a week.
- Open the site on a cheap Android over mobile data. Roughly half of what a test would have found is visible in ninety seconds this way.
Why the sequencing is arithmetic, not a sales preference
We sell SEO, so "do SEO first" from us deserves suspicion. Here's the version that doesn't require trusting us.
The required sample is set by the formula and you can't argue it down. That leaves exactly three ways to reach it. Raise the conversion rate — which is the thing you wanted to test in the first place. Accept a far larger detectable effect, which in practice means redesign-scale changes rather than copy and layout. Or increase traffic. At the volumes most Indian sites actually run at, only the third is something anyone can sell you honestly.
That's the whole of it. Not a philosophy about channels, not a claim that conversion work doesn't matter — a constraint that falls out of the same equation whoever runs it. It also sets what we won't invoice for: we don't sell a monthly testing retainer to a site that can't reach a sample inside a quarter, because we'd be charging for a procedure that cannot return a result. The threshold version of this argument, with budget splits, is in CRO vs SEO, and what we actually do below the floor is on the conversion rate optimisation page.
One clarification, because the vocabulary collides. An SEO split test — changing titles across half a template's URLs and comparing clicks in Search Console — is a different procedure with a different unit of analysis and its own sample calculation. Low traffic hurts it too, but everything on this page is about conversion tests measured on users.
And the number we're accountable for was never a conversion rate. It's your trailing-90-day qualified leads from organic search, frozen the day we start. Miss it in 90 days and we keep working free until we beat it. That's a commitment we can only make because we picked the input we can actually move.