Free calculator

A/B test significance calculator

An A/B test result is statistically significant when its p-value is below your significance level — 0.05 for 95% confidence. Enter visitors and conversions for each variant to get the p-value, uplift and confidence interval instantly.

Updated

Variant A (control)
Variant B
Test settings
Confidence level
Hypothesis

Relative uplift (B vs A)

+20%

A converts at 3%, B at 3.6% — a gap of +0.6 pp.

Statistically significant at 95% confidence (p = 0.0175).

p-valueTwo-sided
0.0175
Observed confidence1 − p
98.25%
z-score
2.375
95% interval for B − APercentage points
+0.1 pp to +1.1 pp

Runs entirely in your browser — nothing you type is sent anywhere. The URL updates as you type, so you can bookmark or share the exact numbers.

What does statistical significance mean in an A/B test?

An A/B test shows two versions of a page to two random groups and compares their conversion rates. Because each group is a sample, the rates would differ a little even if both versions were identical. Statistical significance asks: if the versions really performed the same, how surprising would a gap this large be?

That surprise is the p-value. A p-value of 0.018 means that, if there were no real difference, you’d see a gap at least this large about 1.8% of the time. At a 95% confidence level you accept results with p below 0.05, so you call it significant — knowing that roughly 1 in 20 such calls on truly identical variants will be a false alarm.

The formula: two-proportion z-test

pA = conversionsA ÷ visitorsA pB = conversionsB ÷ visitorsBp = (conversionsA + conversionsB) ÷ (visitorsA + visitorsB)SE = √( p × (1 − p) × (1/visitorsA + 1/visitorsB) )z = (pB − pA) ÷ SEp-value (two-sided) = 2 × (1 − Φ(|z|))Uplift = (pB − pA) ÷ pA
Φ is the standard normal CDF. The confidence interval for pB − pA uses the unpooled standard error √(pA(1−pA)/nA + pB(1−pB)/nB).

The normal approximation behind this test works well once each variant has at least a few dozen conversions and non-conversions. With single-digit conversion counts, treat any result as directional and keep collecting data.

Worked example

A pricing page test sends 10,000 visitors to each variant. The control (A) gets 300 signups, the new layout (B) gets 360.

  1. Rates: pA = 300 ÷ 10,000 = 3.0%, pB = 360 ÷ 10,000 = 3.6%. Relative uplift = 0.6 ÷ 3.0 = +20%.
  2. Pooled rate: p = 660 ÷ 20,000 = 3.3%.
  3. Standard error: √(0.033 × 0.967 × (1/10,000 + 1/10,000)) ≈ 0.002526.
  4. z = 0.006 ÷ 0.002526 ≈ 2.375.
  5. Two-sided p-value = 2 × (1 − Φ(2.375)) ≈ 0.0175 — below 0.05, so significant at 95% confidence.
  6. 95% confidence interval for the difference: about +0.10 to +1.10 percentage points. B is very likely better, but the true lift could be anywhere from a few percent to over 35%.

Now change B to 340 conversions in the calculator: p rises to about 0.11 and the result is no longer significant. Small differences in conversions move the verdict a lot at this traffic level, which is why the test length should be fixed up front with the sample size calculator.

How often do A/B tests win? Benchmarks

Expect most ideas not to win. In “Online Experimentation at Microsoft” (Kohavi et al., 2009), the authors report that only about one third of the ideas tested improved the metric they were designed to improve; the rest were flat or negative. Win rates vary by team and product, but that base rate is a useful corrective: a test result that seems too good is more often a statistical fluke or a tracking bug than a breakthrough.

Confidence levelα (significance)Two-sided z thresholdFalse positives on identical variants
90%0.101.645About 1 in 10
95%0.051.960About 1 in 20
99%0.012.576About 1 in 100
95% is the common default; use 99% for changes that are expensive to reverse.

Common A/B testing mistakes

  • Peeking and stopping early. Checking every day and stopping the first time p dips below 0.05 makes false positives far more likely than 5%. Fix the sample size first, then evaluate once.
  • Testing many variants or metrics and reporting the winner. With 5 variants vs control at α = 0.05, the chance that at least one looks significant by luck is about 23%. Correct for it (e.g. Bonferroni: α ÷ number of comparisons).
  • Choosing one-sided after seeing the data. A one-sided test halves the p-value; pick it before the test only if you’d never ship a version that’s worse.
  • Ignoring sample ratio mismatch. If you split 50/50 but got 10,000 vs 9,200 visitors, the assignment or tracking is broken and the result can’t be trusted.
  • Counting bots. Crawlers rarely convert and can land unevenly in variants; filter them before counting visitors.
  • Stopping before a full week. Weekday and weekend visitors behave differently; run whole weeks.

How to get A/B test numbers from VisitTrack

VisitTrack isn’t an experimentation platform — it doesn’t split traffic for you. But if your testing tool, feature flag or own code assigns the variant, it gives you clean visitor and conversion counts for this calculator, with bots already removed.

  1. Serve each variant on its own URL (e.g. /pricing and /pricing-b), or fire an event per variant view such as window.visitrack("pricing_view_b") (custom events).
  2. Create a funnel per variant — variant view → signup. The Funnels tab shows entered, completed and the conversion rate for each.
  3. Copy the entered and completed numbers into the calculator above.
  4. If the test is about revenue rather than signups, connect your payment provider (revenue attribution) and compare revenue per visitor by landing page, not just conversion rate.

Frequently asked questions

How do I know if my A/B test is statistically significant?

Compare the p-value with your significance level. Enter visitors and conversions for both variants; if the p-value is below 0.05 (for 95% confidence) the difference is statistically significant. Above it, the gap could plausibly be random noise.

What p-value is significant for an A/B test?

A p-value below 0.05 is the most common threshold, matching 95% confidence. Stricter teams use 0.01 (99%) for risky changes, and some accept 0.10 (90%) for cheap, reversible ones. Choose the threshold before the test starts.

What test does this A/B significance calculator use?

It uses a pooled two-proportion z-test, the standard frequentist test for comparing two conversion rates. The confidence interval for the difference uses the unpooled standard error. Bayesian or sequential calculators can give different numbers because they answer slightly different questions.

Should I use a one-sided or two-sided test?

Use two-sided unless you decided in advance that you only care whether B is better and would never act on B being worse. Two-sided is the safer default because it also catches variants that hurt conversion.

Can I stop an A/B test as soon as it reaches significance?

No — not with a fixed-sample test like this one. Repeatedly checking and stopping at the first significant result inflates the false-positive rate well above your chosen α. Calculate the required sample size first and evaluate when you reach it.

What does 95% confidence mean in A/B testing?

It means that if the two variants were truly identical, a test like this would wrongly declare a winner only about 5% of the time. It does not mean there’s a 95% chance B is better.

Why does a 20% uplift sometimes come out as not significant?

Because uplift says nothing about sample size. A jump from 3% to 3.6% is significant with 10,000 visitors per variant (p ≈ 0.018), but the same rates with 2,000 visitors per variant give p ≈ 0.29. Small samples produce big, unreliable swings.

Related tools and guides

Stop calculating by hand — measure it automatically

VisitTrack tracks visitors, goals, funnels and revenue from Stripe, Paddle, Polar, Lemon Squeezy and Razorpay, so these numbers are always on your dashboard. Cookie-free, one script tag, 14 days free.