Answered straight

How much traffic do you need before an A/B test means anything?

The short answer

Count conversions, not sessions. Detecting a 20% relative improvement at 95% confidence and 80% power takes roughly 400 conversions per variant — about 800 in total, at almost any baseline rate. At a 2% conversion rate that's 40,000 sessions. Most Indian sites we quote for collect that in a year, not a month.

Updated 2 October 2026 · Written by the Last Agency team · See what SEO actually costs

The short version

  • Count conversions, not sessions. A +20% test needs roughly 400 conversions per variant whatever your conversion rate. Sessions are just that number divided by the rate.
  • Three inputs decide everything: your baseline conversion rate on that specific template, the smallest lift you'd act on, and how long you'll wait.
  • Halving the effect you want to detect roughly quadruples the sample. That one fact is why "we just want to know if it's better" is an unanswerable request.
  • A higher conversion rate doesn't reduce the conversions you need. It reduces the sessions it takes to collect them.
  • If the arithmetic says no, stop calling the work a test. Don't run one anyway and read the dashboard.

Three inputs decide it, and traffic is the output

An A/B test is a statistical procedure with a fixed appetite. Feed it less than it needs and it doesn't return a smaller answer — it returns an unreliable one, which is worse, because on screen an unreliable answer looks exactly like a reliable one.

Two of the four inputs are conventions you shouldn't be negotiating: 95% two-sided significance, meaning a 5% chance of calling a difference that isn't there, and 80% power, meaning a 20% chance of missing one that is. The other two are yours, and the second is where most requests fall apart.

  • Your baseline conversion rate on the specific page or template you'd test — not the site average. A blog post converting at 0.3% and a pricing page at 6% are different experiments with wildly different appetites.
  • The minimum detectable effect: the smallest improvement you'd actually act on, stated as a relative lift. "Any improvement" implies an infinite sample. Somebody has to name a number, and it should be the person who'd have to justify the change.
  • How long you'll let it run. Weekly traffic times weeks is the entire budget. Two weeks minimum so you capture a full weekly cycle, and rarely beyond six before seasonality starts contaminating the comparison.

The number that matters is conversions, not sessions

Almost every discussion of this starts with traffic, which is why it stays confusing. Start with conversions instead and the arithmetic gets memorable.

  • The conversions column barely moves with your conversion rate. Run it at 1% or at 6% and the figure stays within a few percent of what's above. A better-converting page doesn't need fewer conversions — it collects them from fewer sessions. That's the whole reason to hold conversions in your head instead of traffic.
  • Sample scales with the square of the effect. Detecting +10% instead of +20% doesn't double the requirement, it roughly quadruples it. This is the most expensive decision in the exercise and it usually gets made in passing.
  • Most conversion improvements that survive contact with reality land between 5% and 20% relative. Read the top three rows again with that in mind.
Sample required per variant at 95% two-sided significance and 80% power.
Relative lift you want to detectConversions per variantSessions per variant at 2%Total sessions for the test
+5%about 6,200about 315,000about 630,000
+10%about 1,600about 81,000about 161,000
+20%about 415about 21,000about 42,000
+30%about 190about 9,800about 20,000
+50%about 75about 3,800about 7,600

Four ordinary setups, run through the same formula

The businesses below are illustrative, not clients. The arithmetic is real, and the pattern in it is the point of the page.

  1. Rows one and two: don't test. Not "test carefully" or "test for longer" — don't. A twelve-month experiment isn't an experiment, because the site, the market and the traffic mix all change underneath it.
  2. Row three is the surprise. It has the second-smallest traffic figure in the table and it's the only mid-size site that can test, because a 6% baseline collects the required conversions from far fewer sessions. High-intent B2B pages can start testing much earlier than blog traffic can.
  3. Row four is the argument. Same business, same page, same conversion rate, ten times the traffic — twelve months becomes five weeks. Nothing about the test changed. Only the input that SEO produces.
Time to complete a test detecting a 20% relative lift, at realistic Indian traffic levels.
SetupBaselineMonthly sessions on that templateSessions needed in totalTime to finish
D2C skincare brand, product page1.2%6,000about 71,000about 12 months
Interior design studio, enquiry form2.5%2,000about 34,000about 17 months
B2B software, pricing page6%4,000about 13,000about 3 months
The same D2C brand, four years later1.2%60,000about 71,000about 5 weeks

What calling it early actually does

Every testing platform shows a leading variant and a confidence figure from around day three. The arithmetic is correct and the conclusion is worthless, because the 5% error rate you signed up for assumes you looked once, at the end.

Check daily for a fortnight and you've taken fourteen chances to catch a 5% event. Your real false-positive rate is no longer 5%. How far above depends on how often you look and how quickly you stop, and it is not a small correction — it's the entire reason the sequential-testing literature exists, and why researchers built always-valid p-values for people who can't stop watching.

The failure mode has a recognisable shape. A test is stopped on day four with the variant 18% ahead. It ships. The metric drifts back to where it started over the following six weeks and nobody notices, because there's a new test running and the comparison window has closed. Worse, the phantom win hardens into a belief — "our audience responds to urgency" — and the next three tests get built on top of it.

  1. Write the required sample and the end date down before launch, in the same document as the hypothesis. If it isn't written before, it will be negotiated after.
  2. Run whole weeks. Stopping mid-week weights the sample towards whichever days happened to be included, and weekday and weekend buyers behave differently on almost every Indian site we've looked at.
  3. Name one primary metric in advance. Reading five and reporting the one that moved is peeking with extra steps.
  4. If you genuinely must monitor continuously, use a platform doing sequential or always-valid inference, and accept that it asks for more sample in exchange for letting you look.

What to do when the arithmetic says no

The wrong response is to run the test anyway and describe the outcome as directional. There's no such thing as a directional result from an underpowered test — there's a number, and it's roughly as likely to point the wrong way as the right one.

The right response is to stop using the word test and go and get information some other way. None of the following produces evidence in the statistical sense, and all of it produces better decisions than noise does.

  1. Five customer calls. People who bought last month, fifteen minutes each. Ask what nearly stopped them. Two of the five will say the same thing, and that's your next change.
  2. The objection log your sales team keeps in their heads. Write it down for a fortnight. The top three objections belong on the page, answered in the customer's words.
  3. Session replay on abandonment only. Watch twenty sessions that reached the form and left. You're hunting breakage, not inspiration.
  4. Internal site search queries. People typing into your search box are telling you exactly what the navigation failed to show them — how to read that data is a page of its own.
  5. One exit question. "What stopped you today?", free text. Fifty answers is plenty and it takes a week.
  6. Open the site on a cheap Android over mobile data. Roughly half of what a test would have found is visible in ninety seconds this way.

Why the sequencing is arithmetic, not a sales preference

We sell SEO, so "do SEO first" from us deserves suspicion. Here's the version that doesn't require trusting us.

The required sample is set by the formula and you can't argue it down. That leaves exactly three ways to reach it. Raise the conversion rate — which is the thing you wanted to test in the first place. Accept a far larger detectable effect, which in practice means redesign-scale changes rather than copy and layout. Or increase traffic. At the volumes most Indian sites actually run at, only the third is something anyone can sell you honestly.

That's the whole of it. Not a philosophy about channels, not a claim that conversion work doesn't matter — a constraint that falls out of the same equation whoever runs it. It also sets what we won't invoice for: we don't sell a monthly testing retainer to a site that can't reach a sample inside a quarter, because we'd be charging for a procedure that cannot return a result. The threshold version of this argument, with budget splits, is in CRO vs SEO, and what we actually do below the floor is on the conversion rate optimisation page.

One clarification, because the vocabulary collides. An SEO split test — changing titles across half a template's URLs and comparing clicks in Search Console — is a different procedure with a different unit of analysis and its own sample calculation. Low traffic hurts it too, but everything on this page is about conversion tests measured on users.

And the number we're accountable for was never a conversion rate. It's your trailing-90-day qualified leads from organic search, frozen the day we start. Miss it in 90 days and we keep working free until we beat it. That's a commitment we can only make because we picked the input we can actually move.

Sources

  1. Always Valid Inference: Bringing Sequential Analysis to A/B TestingarXiv · 2015-12-15
  2. About Analytics sessionsGoogle Analytics Help

Every source above was checked on 2 October 2026.

Related questions.

How many conversions do I need per variant for an A/B test?

Roughly 400 to detect a 20% relative lift, about 1,600 for a 10% lift, and about 190 for a 30% lift — at 95% confidence and 80% power. That figure stays nearly constant across baseline conversion rates, which is why conversions are a more useful unit than sessions.

Can I run a valid A/B test with 1,000 visitors a month?

Not for anything realistic. At 1,000 sessions and a 2% conversion rate you collect 20 conversions a month, and a 20% lift needs roughly 840 across both variants — about three and a half years. The only effects detectable at that volume are so large you'd ship them without testing.

Does a higher conversion rate mean I need less traffic?

Yes, and roughly in proportion — doubling your conversion rate halves the sessions required. It does not reduce the number of conversions you need, which is why a high-intent B2B page on 4,000 monthly sessions can test while a blog on 20,000 cannot.

Can I just run the test for longer instead?

Up to a point. Below two weeks you haven't seen a full weekly cycle. Beyond about six, seasonality, cookie churn and every other change you shipped meanwhile start contaminating it, and you're comparing two different periods rather than two variants.

Are 95% confidence and 80% power negotiable?

Technically yes, practically no. Dropping to 90% confidence and 70% power cuts the sample by around 40%, doubles your false-positive rate to 10%, and raises the chance of missing a real win from 20% to 30%. You've bought speed by making the result mean less.

Keep reading

Next, the thing you’ll ask after this.

Last slot's open

Make this the last growth call you book.

Grab the free strategy call and walk away with a 90-day growth plan — hired or not. Or just text us. Either way, you'll know exactly how we'd win.

Guaranteed or it's free · No lock-in · Free strategy call