Free calculator

A/B test sample size calculator

An A/B test needs enough visitors per variant to detect your minimum detectable effect at the significance and power you choose — 13,914 per variant to catch 3% → 3.6% at 95% confidence and 80% power. Enter your numbers to get yours and the test duration.

Updated

What you expect
%

The control’s current rate.

%

Smallest change worth detecting.

Effect is
Statistics
%

5% = 95% confidence.

%

80% is the usual default.

Hypothesis
Traffic

All variants together — used to estimate how long the test runs.

Visitors needed per variant

13,914

To detect 3% → 3.6% reliably.

Total visitors2 variants
27,828
Expected rate in the variant
3.6%
Test duration≈ 2.7 weeks
19 days

Runs entirely in your browser — nothing you type is sent anywhere. The URL updates as you type, so you can bookmark or share the exact numbers.

What does an A/B test sample size calculator tell you?

It tells you how many visitors each variant needs before the test can reliably detect the smallest improvement you care about. Too few visitors and a real improvement looks like noise (a false negative); stop early and noise looks like an improvement (a false positive). Choosing the sample size up front is what keeps both error rates at the levels you picked.

Four inputs decide it: your baseline conversion rate, the minimum detectable effect (MDE), the significance level α (how often you accept a false positive), and statistical power (how often you catch a real effect of that size). Daily traffic then turns the answer into a test duration.

Sample size formula for comparing two conversion rates

p1 = baseline rate p2 = p1 × (1 + relative MDE)p̄ = (p1 + p2) ÷ 2n = ( z(1−α/2) × √(2·p̄·(1−p̄)) + z(1−β) × √(p1(1−p1) + p2(1−p2)) )² ÷ (p2 − p1)²Total = n × number of variants Days = Total ÷ daily visitors
n is visitors per variant, rounded up. z(1−α/2) = 1.96 for 95% two-sided; z(1−β) = 0.84 for 80% power. One-sided tests use z(1−α).

This is the standard normal-approximation formula for a two-sample test of proportions without continuity correction. Other calculators may differ by a few percent — some add a continuity correction, use an arcsine transformation, or run sequential tests that allow early stopping — but they land in the same range.

Worked example

  1. Your signup page converts at 3%. You only care about changes of +20% or more, so p2 = 3% × 1.2 = 3.6%.
  2. With α = 5% (two-sided) and 80% power: z = 1.960 and 0.842, and p̄ = 3.3%.
  3. First term: 1.960 × √(2 × 0.033 × 0.967) ≈ 0.4951. Second term: 0.842 × √(0.03 × 0.97 + 0.036 × 0.964) ≈ 0.2126.
  4. n = (0.4951 + 0.2126)² ÷ 0.006² ≈ 13,913.6 → 13,914 visitors per variant, 27,828 total.
  5. At 1,500 visitors a day entering the test, that’s 19 days — round up to three full weeks.

How many visitors does an A/B test need? Reference table

Baseline rateRelative MDEPer variant (80% power)Per variant (90% power)
1%+20%42,69357,154
2%+20%21,10928,258
3%+10%53,21171,233
3%+20%13,91418,626
3%+30%6,4558,640
5%+10%31,23441,813
5%+20%8,15810,921
10%+10%14,75119,747
10%+20%3,8415,142
Computed with the formula above at α = 0.05, two-sided. Two variants (A/B); multiply by the number of variants for the total.

Which significance and power should you use?

The common defaults are α = 0.05 and 80% power. The 80% convention comes from Jacob Cohen’s Statistical Power Analysis for the Behavioral Sciences (2nd ed., 1988), which proposed 0.80 as a reasonable default when there’s no better basis for choosing. Raise power to 90% when missing a real improvement is costly; lower α to 0.01 when shipping a false winner is costly. Both choices increase the sample size, as the table shows.

Common sample size mistakes

  • Choosing an MDE from wishful thinking. Most changes move conversion by a few percent, not 30%. An MDE you can afford to detect with your traffic is more useful than one you hope for.
  • Confusing relative and absolute effects. +20% relative on 3% is 3.6%; +20 points absolute would be 23%. The calculator lets you pick either.
  • Running several variants without adjusting. Each extra variant adds traffic and, unless you correct α, extra chances of a false winner.
  • Ending at the sample size mid-week. If the sample fills on a Tuesday, keep going to the end of the week so every weekday is represented equally.
  • Using total site traffic. Only visitors who actually reach the tested page count. A checkout test gets checkout traffic, not homepage traffic.
  • Peeking. A fixed sample size only protects you if you evaluate once, at the end.

How to get your baseline and traffic from VisitTrack

  1. Find the page’s daily visitors on the Pages report, with bots already filtered out — the number to put in the traffic field.
  2. Track the conversion with one line, window.visitrack("signup"), or a data-vt-goal attribute (custom events), and add it as a goal.
  3. Build a funnel from the tested page to the conversion: its conversion rate over the last 30 days is your baseline.
  4. Pull the same numbers from a script or an AI assistant via the REST API and MCP server if you plan tests regularly.

Frequently asked questions

How do I calculate sample size for an A/B test?

Use your baseline conversion rate, the minimum detectable effect, the significance level and the power in the two-proportion sample size formula. For a 3% baseline, a +20% relative effect, α = 0.05 and 80% power, that gives 13,914 visitors per variant.

What is a minimum detectable effect (MDE)?

The MDE is the smallest change in conversion rate the test is designed to detect reliably. Smaller MDEs need far more traffic — sample size grows roughly with 1 ÷ MDE², so halving the MDE about quadruples the visitors needed.

What power should an A/B test have?

80% is the standard default, following Jacob Cohen’s 1988 recommendation. It means the test catches a real effect of the MDE size four times out of five. Use 90% if missing a true winner is expensive; it needs about a third more traffic.

How long should an A/B test run?

Until each variant reaches the planned sample size, rounded up to whole weeks. Divide the total sample by the daily visitors entering the test, and run at least one full week even if the sample fills sooner.

Why does a low conversion rate need a bigger sample?

Because rare events are noisy. At a 1% baseline, a +20% lift means just 0.2 extra conversions per 100 visitors, so you need about 42,700 visitors per variant to see it reliably, versus 3,841 at a 10% baseline.

Is the sample size per variant or total?

The main result is per variant. Multiply it by the number of variants, including the control, for the total traffic the test needs. The calculator shows both.

Why do different sample size calculators give slightly different numbers?

They use different approximations. Some add a continuity correction, some use an arcsine transformation, and sequential or Bayesian tools answer a different question. Differences of a few percent are normal; differences of 2x usually mean relative vs absolute MDE got mixed up.

Related tools and guides

Stop calculating by hand — measure it automatically

VisitTrack tracks visitors, goals, funnels and revenue from Stripe, Paddle, Polar, Lemon Squeezy and Razorpay, so these numbers are always on your dashboard. Cookie-free, one script tag, 14 days free.