Tools marked “Coming soon” are being built now.

Business tools
Business · Free · Private

A/B Test Calculator

Type your visitors and conversions. Get a plain-English verdict, the full statistics, and how long to run the next test.

  • Runs in your browser
  • No upload
  • No sign-up
Confidence level

Your test numbers

—
—

Result

—

B against A

Conversion rates
—
Relative uplift
—
p-value (two-sided)
—
Difference, 95% interval
—
Chance it beats A
—
z-score
—

Decide the sample size first, then look once. Checking every day and stopping the moment it says “significant” makes false winners far more likely than the confidence level suggests.

Run whole weeks so weekday and weekend visitors are both counted, and treat a surprisingly huge lift with suspicion until it repeats.

About this tool

A/B Test Calculator tells you whether the difference between two versions of a page, email or advert is real or just the noise you would expect from random visitors. Type the visitors and conversions for the original and the new version, and it shows each conversion rate, the relative uplift, a two-sided p-value from a two-proportion z-test, the confidence interval of the difference, and a verdict in ordinary words, such as “Not enough evidence yet”.

It also gives the Bayesian view many people find easier to act on: the chance that B is truly better than A, worked out exactly rather than estimated. You can add a third and fourth variant; the calculator then shows each one against the control, the chance that each is the best overall, and can split the significance level between the comparisons so extra variants don’t produce extra false winners. A quick check warns you when visitors were split far more unevenly than chance allows, a common sign of a broken test set-up.

The second tab plans tests before you start: from your current rate and the smallest change worth finding, it works out the visitors needed per variant and how many days that takes at your traffic. Everything is calculated in your browser.

How to use A/B Test Calculator

  1. In “Check a result”, type the visitors and conversions for the control (A) and the variant (B). Rename them if you like.
  2. Pick a confidence level: 90%, 95% or 99%. 95% is the usual choice.
  3. Read the verdict, then the detail card: rates, uplift, p-value, interval and the chance B beats A.
  4. Testing more than two versions? Press “Add a variant” for C and D. Press “Copy summary” to paste the result into a report.
  5. Before your next test, open “Plan a test”, enter your current rate, the smallest change worth finding and your daily visitors.
Example

5,000 visitors and 200 sales on A (4%) against 5,000 visitors and 230 sales on B (4.6%) is a 15% uplift, but the p-value is 0.139 and the 95% interval runs from −0.2 to +1.4 percentage points, so the verdict is “Not enough evidence yet”. There is still a 93% chance B is better; to confirm a lift that size you would need about 18,000 visitors per variant.

Features

  • Two-proportion z-test with a two-sided p-value and the z-score.
  • Confidence interval of the difference in percentage points at 90%, 95% or 99%, drawn as a bar against zero.
  • Exact Bayesian probability that each variant beats the control, using flat Beta priors.
  • Up to four variants, with a Bonferroni adjustment and a simulated chance that each one is best.
  • Visitor-split check (sample ratio mismatch) using a chi-square test.
  • Sample size planner with relative or absolute change, 80% or 90% power, and two to four variants.
  • Test length in days and whole weeks at your daily traffic, plus the smallest lift a given sample can detect.
  • A copyable text summary with every figure and the verdict.

Tips and good to know

  • Decide the sample size before you start and judge the result once, at the end. Stopping as soon as the p-value dips under 0.05 can make false winners several times more likely.
  • Run tests in whole weeks. Monday visitors often behave differently from Saturday visitors, and a three-day test can catch only one kind.
  • “Not significant” does not mean “no difference”. It means this sample can’t tell; the planner shows the smallest change it could have found.
  • Count visitors, not page views, and count each visitor once. Repeated visits from the same person break the maths behind every A/B calculator.
  • If a small change shows a huge lift, check the tracking before celebrating. Real lifts above 30% are rare on mature pages.

Frequently asked questions

Are my test numbers uploaded anywhere?

No. Every calculation runs in your browser on your own device. Nothing you type is sent anywhere or stored.

Is it free? Are there limits on visitors or variants?

It is free with no sign-up. You can enter up to four variants and visitor numbers in the millions; the exact Bayesian figure switches to a precise approximation only for very large conversion counts.

Does it work on a phone or offline?

Yes. The layout fits phone screens, and once the page has loaded all the maths keeps working without an internet connection.

What does the p-value actually mean?

It is the chance of seeing a gap at least this big between A and B if the change really made no difference. A p-value of 0.03 means luck alone would produce a gap like yours about 3% of the time. It is not the chance that B is better.

Why do the p-value and the chance B beats A disagree?

They answer different questions. The p-value asks how surprising your data would be if nothing changed; the Bayesian figure asks how likely B is to be better given your data. A 93% chance B is better can sit alongside a p-value of 0.14, which is why the verdict uses the stricter test.

Which confidence level should I choose?

95% is the common standard. Use 90% for cheap, easy-to-reverse changes where speed matters, and 99% for expensive or risky changes such as pricing, where a false winner would hurt.

What is the Bonferroni adjustment for?

Each extra variant is another chance for luck to produce a false winner. The adjustment divides the allowed error between the comparisons, so three variants at 95% are each tested at about 98.3%.

Page last reviewed