How long the test runsbefore you can believe it.

Calling a test early is how teams ship losers and kill winners. Enter your baseline rate, the lift worth detecting, and your daily traffic — this works out the sample size each variant needs and the exact number of days you'll wait for an answer you can actually trust.

Try

Your numbers

%

Your current control rate — the page or flow you're testing against.

%

The smallest RELATIVE improvement worth catching, e.g. 10 means a 10% lift (5.0% → 5.5%).

Total traffic per day entering the test, split across every variant.

Control plus each challenger. 2 = a classic A/B; 3+ is an A/B/n test.

The verdict

Slow — over a month of waiting

Days to run

32

Sample / variant

31,234

Total sample

62,468

Detectable variant rate

5.5%

The rate the variant must hit for you to call a win.

That's 5 full weeks. Run it for 35days so every weekday is represented equally and stopping mid-week doesn't bias the result.

Timeline at your traffic

Days to run
≤2 wks>8 wks

Sensitivity · the lift you'll accept sets the calendar

Days to run by detectable lift

Detectable liftSample / variantDays to run
5%122,125123
10%31,23432
15%14,19315
20%8,1589

Smaller wins are expensive. Dropping from a 20% to a 5% detectable lift roughly quadruples the sample — and the days — because you're asking the test to resolve a difference four times finer. Pick the smallest lift that would actually change what you ship, not the smallest one you can imagine.

Pairs with

Build the full test plan

This calculator answers how long. Decide the smallest lift worth chasing with the minimum detectable effect calculator, confirm the raw counts with the A/B test sample size calculator, and once the test is done, check the result is real with the A/B test significance calculator. Together they take you from plan to verdict without a guess in the middle.

Your move

The math gives you a date. I make the test worth running.

Run this A/B test 32 days (31,234 visitors per variant) to detect a 10% lift on a 5% baseline at 95% confidence.

A perfectly powered test on a weak idea still tells you nothing useful. Bring me the page or funnel you're about to test and I'll help you pick the change that's actually worth your traffic — the bold swing that earns a real lift instead of a month spent confirming a tie. Free call, and the test plan is yours either way.

Plain English

Test duration is sample size in disguise.

An A/B test isn't done when it looks done. It's done when each variant has collected enough conversions that the difference you're seeing is unlikely to be noise. That required count is your sample size, and dividing it by your daily traffic is what turns into a number of days.

Two forces set the sample size: how rare your conversion is (a 2% baseline needs far more visitors than a 20% one) and how small a change you're trying to spot. Chasing a tiny 5% relative lift can need four times the traffic of a comfortable 20% lift. That trade-off is the whole game — decide the smallest win worth waiting for, and the math tells you the price in days.

This calculator uses the standard two-proportion power formula behind every serious testing platform. Set your confidence (how sure you want to be the effect is real) and power (your odds of catching a real effect if it exists), and it returns the sample per variant, the total sample, and the days to run at your traffic level — so you commit to a stop date before you start, not after the numbers start tempting you.

The formula

n per variant = ((z_α·√(2·p̄·(1−p̄)) + z_β·√(p₁·(1−p₁) + p₂·(1−p₂)))²) ÷ (p₂ − p₁)²

Baseline 5% (p₁ = 0.05), a 10% relative lift target (p₂ = 0.055), 95% confidence (z_α = 1.96) and 80% power (z_β = 0.8416). That works out to roughly 31,000 visitors per variant — about 62,000 total for a 2-way test. At 2,000 visitors a day, you're running it 31 days. Halve the lift you'll accept and the days roughly quadruple.

Rough sample sizes before you trust the date

The formula gives you a floor, not a calendar. Here's how the price changes with the lift you'll accept — rough sample per variant for a 5% baseline at 95% confidence and 80% power, so you can sanity-check the calculator above against the shape of the trade-off:

Detectable liftSample / variantWhat it means
5%~122,000Resolving a sliver of a difference — four-plus weeks at decent traffic.
10%~31,000The common default. About a month at 2,000 visitors/day on a 2-way test.
15%~14,000A solid, actionable win — readable in a couple of weeks at moderate traffic.
20%~8,000A confident swing. Cheap to confirm; this is the lift bold changes earn.

Rule-of-thumb sample sizes — the power formula at a 5% baseline, not a published dataset. Halving the lift you'll accept roughly quadruples the sample. Use the calculator above for your own numbers.

Your test wants to run forever. Now what?

01

Duration is months, not weeks.

You're either chasing too small a lift or you don't have the traffic. Raise the minimum detectable lift to the smallest win that would actually change a decision, or test a bolder, higher-impact change instead of a tweak.

02

Baseline conversion is very low (under ~2%).

Rare events need huge samples. Move the test up the funnel to a higher-frequency metric (add-to-cart, signup-started) that correlates with the money, then validate the downstream effect separately.

03

You added a third or fourth variant.

Every extra arm splits your traffic and stretches the timeline. Each variant still needs its own full sample — an A/B/C/D doesn't cost the same as an A/B. Cut to your two strongest ideas unless traffic is plentiful.

04

The date lands mid-week.

Always run in whole weeks. Buyer behaviour swings hard between weekdays and weekends; stopping on a Wednesday bakes a day-of-week bias into the result. Round up to the next full 7-day cycle.

How to get a trustworthy answer faster

01Test bigger swings

A radical redesign produces a larger lift than a button-colour tweak, and larger effects need far less traffic to confirm. Bold changes are cheaper to validate, not riskier.

02Pick the lift you'd act on

Don't set the minimum detectable lift to the smallest number you can imagine. Set it to the smallest win that would actually change what you ship. That single choice controls the whole timeline.

03Raise the baseline you test

Run the experiment on a higher-converting metric or a higher-intent audience. The rarer the event, the more traffic it takes to read.

04Concentrate traffic

Funnel more of your audience into the test, or test on your highest-traffic page first. Fewer simultaneous tests means each one finishes sooner.

05Commit to the stop date up front

Decide the sample and date before launch, then don't peek-and-stop. Repeatedly checking and calling it the moment it looks significant inflates false positives badly.

06Run full weeks, every time

Always complete whole 7-day cycles so every weekday is represented equally. A test that ends mid-week is silently biased by which days it captured.

07Limit the number of variants

Two strong ideas beat four weak ones. Each arm needs its own full sample, so every variant you add buys you more waiting.

08Don't undersize to ship faster

Stopping at half the required sample doesn't get you a faster answer — it gets you a confident wrong one. The duration is the price of the certainty.

The vocabulary

Sample size
The number of visitors each variant must collect before the result is statistically reliable. Total sample = per-variant × number of variants.
Baseline conversion rate
The conversion rate of your control. Lower baselines need much larger samples to detect the same relative change.
Minimum detectable effect (MDE)
The smallest relative lift the test is powered to catch. A 10% MDE on a 5% baseline means detecting a move to 5.5%.
Statistical confidence
1 − α: how sure you want to be that a detected effect is real, not chance. 95% is the common default; the z-value is 1.96.
Statistical power
1 − β: the probability of detecting a real effect when one truly exists. 80% is standard; the z-value is 0.8416.
Two-sided test
A test that checks whether the variant is better OR worse than control, not just better. It's the honest default and uses the larger z-value.

Test duration questions, straight answers

Long enough to collect the required sample size at your traffic level, and never less than one full week. For a 5% baseline chasing a 10% relative lift at 95% confidence and 80% power, that's roughly 31,000 visitors per variant — about 31 days at 2,000 visitors a day. Lower traffic or a smaller target lift pushes it longer. Enter your own numbers above to get the exact date.

A calculator tells you what. A call tells you what to do about it.

Send me the account behind these numbers. I'll tell you straight where the money's leaking and what I'd fix first — free, and you keep it whether you hire me or not.