Tools · Decision-first
A/B Test Calculator
Two numbers most dashboards won’t give you straight: is B actually better than A, and how sure can you be? Enter the results to get lift, a p-value, and a verdict. Then size the next test before you launch it. Rolling out a training instead of a landing page? The bottom half sizes the holdout you need and reads whether the trained reps actually pulled away.
Did B beat A?
—
—
The test: a two-proportion two-tailed z-test. The p-value is the chance of seeing a gap this large if A and B were truly equal; below 0.05 is the usual bar for “real.” The interval is the 95% range for the true difference in rates. It assumes a fixed sample decided in advance — peeking at a live test and stopping when it looks good inflates false positives.
How many do I need?
—
Standard two-proportion power calculation at 95% significance. Halve nothing — that count isper variant, so a two-arm test needs roughly double. Smaller effects cost dramatically more traffic; that trade is the whole planning conversation.
How many holdouts?
For a rollout on people — a sales training, a new call script, a coaching program — the test is the same rep measured before and after, with some reps held back so you can see what would have happened anyway. This sizes that holdout.
—
—
| Holdout | Reps held | Catches |
|---|
—
The design: randomize who waits, measure everyone the same way before and after, then compare the trained reps’ post numbers to the holdout’s after adjusting for where each rep started (ANCOVA, the sharper cousin of a plain before/after difference). A holdout of share h keeps 4·h·(1−h) of the information a 50/50 split would: 20% keeps 64%, 10% keeps 36%. Shrinking the holdout is cheap until it isn’t. The holdout gets the training after the read — nobody is denied it, only sequenced.
Did the training work?
—
—
—
The test: each rep’s change (post − pre) is the unit. The training effect is the trained group’s average change minus the holdout’s — a difference-in-differences — checked with a Welch t-test. Whatever moved both groups (seasonality, a comp-plan change, a ramping cohort) cancels out; that is the whole reason for the holdout. “SD of per-rep change” is one STDEV over the change column. If you have the rep-level rows, regress post on pre and a trained flag instead: same answer, tighter interval.
Measurement before celebration. Knowing whether a number is real — and how much traffic it takes to find out — is the same discipline I bring to an AI feature in production.