Statistical Significance Calculator
Enter visitors and conversions for both variants of your A/B test. You'll get the p-value, the relative lift, and a straight answer on whether the difference clears the usual 95% bar. Below it, a second calculator tells you how many visitors you need before a result can mean anything — the number most people wish they had checked first.
Free, no signup, and everything is calculated in your browser.
p = 0.0055 · 99.4% confidence
Two-tailed pooled z-test. A p-value is the chance of seeing a difference this large if the two variants were actually identical — not the chance that B is better.
How long to run it
Sample size and test duration for the lift you hope to detect, at 95% confidence and 80% power.
24,008 visitors in total · 12,004 per variant · rounded up to whole weeks
How significance is calculated
A pooled two-proportion z-test. The pooled rate p̂ = (x₁ + x₂) ÷ (n₁ + n₂), the standard error is √(p̂(1−p̂)(1/n₁ + 1/n₂)), and z is the difference in rates divided by that standard error.
The p-value is two-tailed: 2(1 − Φ(|z|)). One-tailed tests halve the p-value and flatter whichever variant happens to be ahead, so they're not offered here.
A p-value is the probability of a difference at least this large if the two variants were genuinely identical. It is not the probability that B beats A.
Required sample size per variant (the duration calculator) uses the standard two-proportion formula: n = (z₁₋α/₂ √(2p̄(1−p̄)) + z₁₋β √(p₁(1−p₁) + p₂(1−p₂)))² ÷ (p₂ − p₁)², with α = 0.05 (95% confidence) and β = 0.2 (80% power). Days = n × 2 ÷ daily visitors, rounded up to whole weeks so weekday and weekend behaviour are both represented.
How is statistical significance calculated for an A/B test?
This calculator runs a pooled two-proportion z-test, the standard large-sample test for comparing two conversion rates (NIST's engineering statistics handbook describes the same test for comparing two defect rates). It asks one question: if A and B truly converted at the same rate, how often would random assignment alone produce a gap as large as the one you saw?
The steps are: work out each variant's rate, pool both groups into a single rate that assumes no difference, use the pooled rate to find the standard error of the difference, then divide the observed difference by that standard error to get z. The p-value is the probability of a z at least that far from zero, in either direction.
Worked example: is 15% really better than 12%?
Illustration, using the calculator's default numbers. Variant A: 2,000 visitors, 240 conversions, so 12.00%. Variant B: 2,000 visitors, 300 conversions, so 15.00%. Relative lift is (15 − 12) ÷ 12 = +25%.
Pooled rate: (240 + 300) ÷ (2,000 + 2,000) = 540 ÷ 4,000 = 13.5%. Standard error: √(0.135 × 0.865 × (1/2,000 + 1/2,000)) = √(0.116775 × 0.001) = 0.01081. z = (0.15 − 0.12) ÷ 0.01081 = 2.78. Two-tailed p = 2 × (1 − Φ(2.78)) ≈ 0.0055, well under 0.05, so the calculator reports Significant.
Now keep the same two rates but shrink the test to 200 visitors each (24 vs 30 conversions). The standard error grows to 0.0342, z falls to 0.88 and p rises to 0.38. Same 25% lift, no evidence at all. Sample size, not the size of the gap, is usually what decides the verdict.
One more: 10,000 visitors each, 1,200 vs 1,320 conversions (12.0% vs 13.2%, a 10% lift). z = 2.56, p ≈ 0.011. Significant, but only because the sample is large; a 10% lift is small and needs a lot of traffic to see.
What z-score and p-value mean significant?
The calculator reports a two-tailed p-value. These are the usual cut-offs and the |z| each one needs.
| |z| at least | Two-tailed p | Shown as confidence | Typical use |
|---|---|---|---|
| 1.645 | 0.10 | 90% | Low-stakes changes; about 1 in 10 null tests will look like a winner |
| 1.96 | 0.05 | 95% | The common default, and the calculator's Significant threshold |
| 2.576 | 0.01 | 99% | Pricing, checkout, anything expensive to reverse |
| 3.29 | 0.001 | 99.9% | Many simultaneous metrics or tests |
The calculator's confidence figure is simply (1 − p) × 100. It is not the probability that B beats A.
Sources: NIST/SEMATECH e-Handbook: comparing two proportions
What does a p-value actually tell you?
A p-value of 0.03 means: if the variants were identical, a difference this large or larger would appear about 3 times in 100 tests. It does not mean there is a 97% chance B is better, and it says nothing about whether the lift is big enough to matter. The American Statistical Association's 2016 statement makes both points and warns against treating 0.05 as a bright line.
In practice, read three numbers together: the p-value (is there evidence of any difference?), the lift (how big is it?), and the sample (was the test sized to detect a lift this big?). A significant 0.3% lift on a pricing page and a non-significant 30% lift on 150 visitors are both results you should not ship on.
Source: ASA Statement on Statistical Significance and p-Values (2016)
One-tailed or two-tailed test?
This calculator is two-tailed only. A one-tailed test counts only differences in the direction you predicted, which halves the p-value: z = 1.80 gives p = 0.072 two-tailed but 0.036 one-tailed, turning Not significant into Significant without a single extra visitor.
One-tailed tests are defensible when a result in the other direction would lead to exactly the same decision as no result, and when the direction is chosen before the data arrive. In product testing you almost always care if B is worse, so two-tailed is the honest default.
Why you shouldn't stop an A/B test the moment it hits significance
The p-value assumes you look once, at a sample size you fixed in advance. Checking every day and stopping at the first significant reading gives chance many extra opportunities to cross the line. In Evan Miller's worked example, a test at a nominal 5% level that is stopped as soon as it looks significant produces false positives 26.1% of the time, more than five times the rate you think you are running.
The fix is dull: decide the sample size first, then do not act on the result until you reach it. If you genuinely need to stop early, use a method built for it, such as a sequential design (Miller publishes a simple one) or the always-valid approaches covered in Kohavi, Tang and Xu's Trustworthy Online Controlled Experiments.
Source: Evan Miller: How Not To Run an A/B Test
Source: Evan Miller: Simple Sequential A/B Testing
Source: Kohavi, Tang & Xu: Trustworthy Online Controlled Experiments (Cambridge University Press)
What sample size do you need for an A/B test?
The second calculator on this page answers that before you start. It needs three inputs: your current conversion rate, the smallest relative lift worth detecting (the minimum detectable effect), and daily visitors across both variants. It fixes significance at 95% and power at 80%, meaning that if the true lift is exactly the size you entered, the test will catch it four times in five.
Illustration with its defaults: a 12% baseline and a 10% relative lift (12% to 13.2%) need 12,004 visitors per variant, 24,008 in total. At 500 visitors a day that is 48 days, which the calculator rounds up to 49 days, seven full weeks, so every weekday is represented equally.
Other calculators may use a different test (unpooled, one-tailed or Bayesian), so the same data can produce a slightly different p-value elsewhere. The order of magnitude is what matters: halving the lift you want to detect roughly quadruples the sample.
A/B test sample size per variant (95% confidence, 80% power)
Computed with the same formula as the duration calculator above. Lift is relative: a 10% lift on a 5% baseline means 5.0% to 5.5%.
| Baseline rate | 5% lift | 10% lift | 20% lift | 30% lift |
|---|---|---|---|---|
| 2% | 315,210 | 80,683 | 21,109 | 9,798 |
| 5% | 122,125 | 31,234 | 8,158 | 3,780 |
| 10% | 57,764 | 14,751 | 3,841 | 1,774 |
| 20% | 25,583 | 6,510 | 1,683 | 772 |
| 30% | 14,856 | 3,763 | 963 | 437 |
Double each figure for the total across both variants. Low-traffic forms can usually only detect large lifts; test bold changes, not button shades.
Testing more than two variants or metrics: the multiple comparisons problem
Every extra comparison is another chance for noise to look like a win. If you run k independent comparisons at the 0.05 level and none of the changes do anything, the chance of at least one false winner is 1 − 0.95^k: 14.3% for 3 comparisons, 22.6% for 5, and 40.1% for 10. Comparing three challengers to a control, or one test across ten metrics, falls into this trap.
Two standard corrections. Bonferroni divides the threshold by the number of comparisons (0.05 ÷ 5 = 0.01 each), which is simple and conservative. Benjamini and Hochberg's 1995 procedure controls the false discovery rate instead and keeps more power when you test many things. Either way, name one primary metric before the test starts.
Source: Benjamini & Hochberg (1995), Controlling the False Discovery Rate, JRSS Series B
When not to trust (or use) this calculator
- Tiny counts: the z-test relies on a normal approximation. With only a handful of conversions in either group, use an exact test such as Fisher's instead.
- Unequal splits you did not plan: if you meant 50/50 and got 55/45, check the randomization before reading the result. Kohavi et al. call this a sample ratio mismatch.
- Revenue or time-on-page: this compares rates (converted or not), not averages. Average order value needs a test for means.
- Visitors counted more than once: if the same person can land in both variants, the groups are not independent.
- Before-and-after comparisons: last month versus this month is not a randomized test. Seasonality and traffic mix change too.
How to A/B test a form and collect clean numbers
For forms, the conversion you usually compare is submissions ÷ visitors who saw the form. Split traffic with your site's testing tool or a redirect that assigns visitors at random, and send each group to a different version of the form.
In Zunoform, the simplest setup is two copies of the form, each with its own link. Each form's Summary tab shows views, starts and completed submissions (the owner's own visits are not counted), which gives you the visitors and conversions to enter above. If both variants must feed one form, declare a hidden field such as variant and put ?variant=b in the B link; the value is stored with each response and exported with it. Zunoform does not randomize the split for you.
Once you know which stage is losing people, our form conversion rate calculator separates start rate from completion rate. Forms worth testing first are usually the high-traffic ones, like a lead capture form, a demo request or a newsletter signup.
Forms people commonly A/B test
Contact details, company size, goals, and buying timeline for inbound leads.
Role, team size, and focus areas so the demo lands on what matters.
Email, topics, and send frequency — a subscribe form people finish.
Pre-launch email capture with use case and urgency for prioritizing invites.
Questions people ask
How do you calculate statistical significance?
For two conversion rates, use a two-proportion z-test. Pool both groups into one rate, compute the standard error √(p(1−p)(1/n₁ + 1/n₂)), divide the difference in rates by it to get z, and convert z to a p-value. If p is below your threshold, usually 0.05, the difference is statistically significant. The calculator above does every step and shows z, p and the lift.
What is a p-value, in plain terms?
The p-value is how surprising your result would be if the two variants were actually identical. A p-value of 0.03 means a gap this large would show up by chance about 3 times in 100. It is not the probability that B is better, and it does not tell you whether the lift is large enough to matter.
How long should I run an A/B test?
Until you reach the sample size you calculated before starting, and never less than one full week, so weekday and weekend behavior are both included. The duration calculator above rounds up to whole weeks for that reason. Stopping the moment a result turns significant is the most common way to ship a change that does nothing.
My result isn't significant. Does that mean the variants are the same?
No. It means the data cannot distinguish them at the sample you collected. Absence of significance is not evidence of equivalence, especially at small samples. Check the sample-size table: if your test was far below the size needed for the lift you care about, it was never able to give a clear answer.
Is 90% confidence good enough?
It depends on the cost of being wrong. At 90% you will call a false winner about one test in ten when nothing really changed. That is acceptable for a headline tweak and reckless for a pricing page. Decide the threshold before the test, not after you have seen which variant it favors.
One-tailed or two-tailed?
Two-tailed, which is what this calculator uses. A one-tailed test assumes you already know which direction the difference goes; it halves the p-value and flatters whichever variant happens to be ahead. Unless a worse result would lead to exactly the same decision as no result, two-tailed is the defensible choice.
Can I use this to compare more than two variants?
Only pair by pair, and each extra comparison raises the chance of a false winner: with five comparisons at 0.05 and no real effects, the chance of at least one false positive is about 23%. Tighten the threshold, for example Bonferroni's 0.05 divided by the number of comparisons, or use a test designed for several groups.
Does this work for email, surveys and ads, not just web pages?
Yes. The maths only needs two randomized groups, each with a count of people and a count of successes: opens against sends, completions against starts, clicks against impressions. Anything that can be expressed that way can be tested here, as long as each person was in only one group.
Further reading
Related tools

Ready when you are
Better forms.
Better data.
Start with a conversation, a document or a single page.
Unlimited forms and 500 responses per month on Free.
No card required.
- No credit card
- Unlimited forms
- 500 responses / month