What a sample size calculation does
A sample size calculation answers one question before you launch: how many visitors does each variant need before the test can tell a real lift from luck? Run a test on too few visitors and the result is a coin flip wearing a lab coat. This calculator uses the standard two-proportion formula for a two-sided test with a 50/50 split.
Here p₁ is your baseline conversion rate, p₂ is the rate your minimum detectable effect (MDE) implies, and p̄ is their average. The z-values come from your settings: zα/2 is 1.96 at 95% significance, and zβ is 0.84 at 80% power. Tighter confidence or higher power raises the z-values, and the sample grows with their square.
Worked example
Take the calculator's defaults: a 3% baseline conversion rate, a 20% relative MDE, 95% significance, 80% power, and 5,000 visitors a day entering the test. The MDE sets the target rate at 3% × 1.20 = 3.6%, an absolute gap of 0.6 points.
Total sample = 2 × 13,914 = 27,828 visitors
27,828 ÷ 5,000 a day ≈ 5.6 days
So this test needs 13,914 visitors per variant, 27,828 total, and finishes in about 5.6 days at that traffic. Change one setting and watch the cost move: 99% significance pushes it to 20,704 per variant, and 90% power pushes it to 18,626. Rigor is never free; the calculator prices it for you.
Relative vs. absolute MDE
The single most common sample size mistake is mixing these up. This calculator uses a relative MDE: a 20% MDE on a 3% baseline means detecting a move from 3% to 3.6%, not from 3% to 23%. In absolute terms that is 0.6 percentage points, which is why the “target rate to beat” card shows 3.6%.
Relative MDEs travel well between pages: a 20% lift means the same thing on a 2% product page and an 8% landing page. Absolute MDEs do not: one point of lift is a 50% improvement on the first and 12.5% on the second. If a tool or teammate quotes an absolute MDE, convert it before comparing: absolute ÷ baseline = relative. Not sure what your baseline actually is? Compute it first with the Conversion Rate Calculator. A wrong baseline makes every number downstream wrong.
The brutal math of small lifts
Sample size scales with roughly the inverse square of the effect you chase. Halve the MDE and the sample roughly quadruples. On the default 3% baseline: a 40% MDE needs 3,782 visitors per variant, the 20% default needs 13,914, and a 10% MDE needs 53,211: over 21 days of the example site's entire traffic, for one test.
The practical read: match your test size to your traffic. A site with 5,000 daily visitors can afford to hunt 15–20% lifts. A site with 500 should be testing swings big enough to clear a 30–50% MDE. That means new offers, new landing page angles, and new hooks, not button colors. Small sites don't have a testing problem; they have a timidity problem. If a change isn't plausibly worth a 30% lift, it isn't worth a test slot at that traffic level.
Pre-committing to n kills the peeking problem
The classic failure mode: launch a test, check the dashboard every morning, and stop the moment it flashes significant. Each peek is another chance for noise to cross the threshold, so a test “peeked” daily can show a false winner several times more often than its stated significance level suggests.
The fix costs nothing: compute n before launch, write it down, and read the result once, when both variants hit the number. That single commitment is what makes the significance level you chose actually mean what it says. When the test does finish, check the result with the A/B Test Significance Calculator. The two tools are the before and after of the same discipline.
Frequently asked questions
What's the difference between relative and absolute MDE?
A relative MDE is a percentage of your baseline: 20% relative on a 3% baseline means detecting 3.6%. An absolute MDE is a fixed gap in percentage points: 0.6 points on 3% is the same test. This calculator takes the relative form because it compares cleanly across pages with different baselines. Divide an absolute MDE by your baseline to convert it.
Why do small lifts need so many visitors?
Conversion data is noisy, and the smaller the true lift, the harder it is to separate from that noise. Sample size grows with the inverse square of the effect: detecting a lift half the size takes roughly four times the visitors. That is why the defaults need 13,914 per variant for a 20% lift but 53,211 for a 10% lift on the same page.
What power should I use?
80% power is the standard default: if the true lift is exactly your MDE, the test catches it four times out of five. Moving to 90% cuts missed winners but raises the default's sample from 13,914 to 18,626 per variant. Use 90% when missing a real winner is expensive, like a pricing or offer test, and 80% for routine iteration.
Should I use 90% or 95% confidence for ad creative tests?
95% is the convention for site tests, where a false winner gets hard-coded into the page. Ad creative decisions are cheaper to reverse, since a losing ad gets paused next week. Many media buyers therefore accept 90%, which drops the default's requirement from 13,914 to 10,960 per variant. Pick by the cost of being wrong, not by habit.
What if I can never reach the sample size?
Don't run the test anyway and squint at the trend; an underpowered test mostly produces false confidence. Instead: raise the MDE by testing bigger changes, test on your highest-traffic page, lengthen the test window, or pool the decision with other evidence and label it a judgment call. A directional read is fine as long as nobody calls it statistics.
Can I stop the test early once it looks significant?
Not if you want the significance level to mean anything. Stopping at the first significant reading inflates false positives well past your chosen threshold, because every peek is another draw from the noise. Commit to n up front and read once. If you genuinely need early stopping, use a sequential testing method designed for it.
Where StefanBrain fits
This calculator tells you what a test costs in traffic; the harder question is what deserves the slot. The brutal math above says most sites can only afford tests that swing big, and big swings are a creative problem. That's where StefanBrain comes in: 7+ years of DTC marketing IP, turned into the angles, landing pages, and copy that justify a 27,828-visitor bet. Sizing paid creative tests instead of site tests? The Creative Testing Budget Calculator translates the same discipline into dollars per concept.
