How much trafficyour test really needs.

Run a test too small and you'll never reach significance. You'll just waste weeks. Enter your baseline conversion rate and the lift you want to catch, set your confidence and power, and get the exact visitors per variant before you launch.

Try

Your test

%

Your current rate for the control.

%

Smallest relative improvement worth catching.

Traffic needed

Detecting 3.0%3.6%

Per variant

13,914

Total visitors

27,828

Target rate

3.6%

Absolute lift

0.6%

Reverse-solve · how long it runs

From traffic to a finish date

Weeks to run

5.6

Days

39

At 5,000 visitors a week, this test needs about 5.6 weeks to reach its sample. Run it in whole weeks so the weekday/weekend mix stays even.

Sensitivity · bolder tests finish faster

Sample size by the lift you target

Lift to detectPer variantTotal
5%207,938415,876
10%53,211106,422
15%24,19348,386
20%13,91427,828
30%6,45512,910
50%2,5185,036

Halving the lift you want to catch roughly quadruples the traffic you need. If a test would take months, the answer is usually to test a bigger swing — not to wait.

Your move

Big enough traffic to test? Let's check before you waste a month.

Detecting a 20% lift on a 3.0% baseline needs 13,914 visitors per variant (27,828 total) at 95% / 80%.

If your traffic can't support the test you're planning, there are faster ways to find the win. Bring me your numbers and I'll tell you what's actually worth testing, and what to just fix. Free, on a 30 minute call.

Plain English

Decide the size before you run the test.

Sample size is the number of visitors each variant needs before a result means anything. It's set by three things: your baseline conversion rate, the smallest lift you want to be able to detect, and how confident and statistically powerful you want the test to be. Skip this step and you're flying blind.

The relationship is unforgiving. Detecting a small lift on a low baseline takes enormous traffic, halving the effect you want to catch roughly quadruples the visitors required. That's why testing a tiny tweak on a low traffic page is often hopeless: the test would need to run for a year to ever conclude.

Calculating it up front protects you twice. It tells you whether a test is even feasible given your traffic, and it forces you to commit to an endpoint so you don't stop early on a fluke. This tool gives you visitors per variant for your inputs, and lets you dial confidence and power.

The formula

n ≈ ( z₁₋α/₂·√[2p̄(1−p̄)] + z₁₋β·√[p₁(1−p₁)+p₂(1−p₂)] )² ÷ (p₂ − p₁)²

Baseline 3%, detect a 20% relative lift (to 3.6%), at 95% confidence and 80% power → roughly 13,900 visitors per variant, about 27,800 total.

What confidence and power mean

Two dials control how rigorous, and how hungry, your test is. The conventional settings, and what moving them costs you:

DialConventionalEffect of raising it
Confidence95%More traffic, fewer false positives
Statistical power80%More traffic, fewer missed winners
Minimum detectable lift10 to 25%Bigger lift = far less traffic needed

Source: Evan Miller / standard two-proportion test conventions · 2025

Read the requirement

01

Required sample is far above your traffic.

The test isn't feasible as scoped. Either aim to detect a bigger lift, or test a higher traffic page where the volume exists.

02

You need a result fast.

Test bolder changes. A larger expected lift needs dramatically less traffic to detect than a marginal tweak.

03

Baseline rate is very low.

Low baselines need huge samples. Consider testing an earlier, higher converting step in the funnel instead.

04

Sample looks reachable in days.

Good, but still run full weeks so the weekday/weekend mix is balanced across both variants.

Make tests finish in this lifetime

01Test bigger swings

Detecting a 30% lift needs a fraction of the traffic of a 5% lift. Bold changes conclude faster and teach you more.

02Test higher up the funnel

Steps with more traffic and higher baseline rates reach significance far sooner than the final purchase.

03Commit to the endpoint

Calculate the sample, then run to it. Pre committing is what makes the eventual significance trustworthy.

04Don't over segment

Splitting traffic into many cells starves each one. Test fewer variants with enough volume each.

05Lower power only deliberately

Dropping from 80% to 70% power shrinks the sample but raises the odds you miss a real winner. Know the trade.

06Run whole weeks

Round the duration up to complete weeks so day of week effects are even across variants.

The vocabulary

Sample size
Visitors needed per variant before the test can detect the target effect reliably.
Baseline conversion rate
The control's current conversion rate, the starting point for the test.
Minimum detectable effect (MDE)
The smallest lift you want the test to be able to catch, here as a relative percentage.
Confidence level
How sure you want to be a detected difference is real. 95% is standard.
Statistical power
The probability of detecting a real effect if one exists. 80% is standard.
Relative vs absolute lift
Relative is the % change (3% → 3.6% is +20%); absolute is the point change (+0.6pp).

Sample size questions

You need three inputs: your baseline conversion rate, the minimum lift you want to detect, and your confidence and power levels (commonly 95% and 80%). Plug them in and the formula returns visitors needed per variant. This tool does it instantly.

A calculator tells you what. A call tells you what to do about it.

Send me the account behind these numbers. I'll tell you straight where the money's leaking and what I'd fix first — free, and you keep it whether you hire me or not.