How much trafficyour test really needs.
Run a test too small and you'll never reach significance. You'll just waste weeks. Enter your baseline conversion rate and the lift you want to catch, set your confidence and power, and get the exact visitors per variant before you launch.
Your test
Your current rate for the control.
Smallest relative improvement worth catching.
Traffic needed
Per variant
13,914
Total visitors
27,828
Target rate
3.6%
Absolute lift
0.6%
Reverse-solve · how long it runs
From traffic to a finish date
Weeks to run
5.6
Days
39
At 5,000 visitors a week, this test needs about 5.6 weeks to reach its sample. Run it in whole weeks so the weekday/weekend mix stays even.
Sensitivity · bolder tests finish faster
Sample size by the lift you target
| Lift to detect | Per variant | Total |
|---|---|---|
| 5% | 207,938 | 415,876 |
| 10% | 53,211 | 106,422 |
| 15% | 24,193 | 48,386 |
| 20% | 13,914 | 27,828 |
| 30% | 6,455 | 12,910 |
| 50% | 2,518 | 5,036 |
Halving the lift you want to catch roughly quadruples the traffic you need. If a test would take months, the answer is usually to test a bigger swing — not to wait.
Your move
Big enough traffic to test? Let's check before you waste a month.
Detecting a 20% lift on a 3.0% baseline needs 13,914 visitors per variant (27,828 total) at 95% / 80%.
If your traffic can't support the test you're planning, there are faster ways to find the win. Bring me your numbers and I'll tell you what's actually worth testing, and what to just fix. Free, on a 30 minute call.
Plain English
Decide the size before you run the test.
Sample size is the number of visitors each variant needs before a result means anything. It's set by three things: your baseline conversion rate, the smallest lift you want to be able to detect, and how confident and statistically powerful you want the test to be. Skip this step and you're flying blind.
The relationship is unforgiving. Detecting a small lift on a low baseline takes enormous traffic, halving the effect you want to catch roughly quadruples the visitors required. That's why testing a tiny tweak on a low traffic page is often hopeless: the test would need to run for a year to ever conclude.
Calculating it up front protects you twice. It tells you whether a test is even feasible given your traffic, and it forces you to commit to an endpoint so you don't stop early on a fluke. This tool gives you visitors per variant for your inputs, and lets you dial confidence and power.
The formula
n ≈ ( z₁₋α/₂·√[2p̄(1−p̄)] + z₁₋β·√[p₁(1−p₁)+p₂(1−p₂)] )² ÷ (p₂ − p₁)²
Baseline 3%, detect a 20% relative lift (to 3.6%), at 95% confidence and 80% power → roughly 13,900 visitors per variant, about 27,800 total.
What confidence and power mean
Two dials control how rigorous, and how hungry, your test is. The conventional settings, and what moving them costs you:
| Dial | Conventional | Effect of raising it |
|---|---|---|
| Confidence | 95% | More traffic, fewer false positives |
| Statistical power | 80% | More traffic, fewer missed winners |
| Minimum detectable lift | 10 to 25% | Bigger lift = far less traffic needed |
Source: Evan Miller / standard two-proportion test conventions · 2025
Read the requirement
Required sample is far above your traffic.
The test isn't feasible as scoped. Either aim to detect a bigger lift, or test a higher traffic page where the volume exists.
You need a result fast.
Test bolder changes. A larger expected lift needs dramatically less traffic to detect than a marginal tweak.
Baseline rate is very low.
Low baselines need huge samples. Consider testing an earlier, higher converting step in the funnel instead.
Sample looks reachable in days.
Good, but still run full weeks so the weekday/weekend mix is balanced across both variants.
Make tests finish in this lifetime
01Test bigger swings
Detecting a 30% lift needs a fraction of the traffic of a 5% lift. Bold changes conclude faster and teach you more.
02Test higher up the funnel
Steps with more traffic and higher baseline rates reach significance far sooner than the final purchase.
03Commit to the endpoint
Calculate the sample, then run to it. Pre committing is what makes the eventual significance trustworthy.
04Don't over segment
Splitting traffic into many cells starves each one. Test fewer variants with enough volume each.
05Lower power only deliberately
Dropping from 80% to 70% power shrinks the sample but raises the odds you miss a real winner. Know the trade.
06Run whole weeks
Round the duration up to complete weeks so day of week effects are even across variants.
The vocabulary
- Sample size
- Visitors needed per variant before the test can detect the target effect reliably.
- Baseline conversion rate
- The control's current conversion rate, the starting point for the test.
- Minimum detectable effect (MDE)
- The smallest lift you want the test to be able to catch, here as a relative percentage.
- Confidence level
- How sure you want to be a detected difference is real. 95% is standard.
- Statistical power
- The probability of detecting a real effect if one exists. 80% is standard.
- Relative vs absolute lift
- Relative is the % change (3% → 3.6% is +20%); absolute is the point change (+0.6pp).
Sample size questions
You need three inputs: your baseline conversion rate, the minimum lift you want to detect, and your confidence and power levels (commonly 95% and 80%). Plug them in and the formula returns visitors needed per variant. This tool does it instantly.
Keep going
A calculator tells you what. A call tells you what to do about it.
Send me the account behind these numbers. I'll tell you straight where the money's leaking and what I'd fix first — free, and you keep it whether you hire me or not.