Tools · Decision-first

A/B Test Calculator

Two numbers most dashboards won’t give you straight: is B actually better than A, and how sure can you be? Enter the results to get lift, a p-value, and a verdict. Then size the next test before you launch it. Rolling out a training instead of a landing page? The bottom half sizes the holdout you need and reads whether the trained reps actually pulled away.

Did B beat A?

Variant A (control)

Variant B (challenger)

—

Rate A—
Rate B—
Relative lift—
p-value—

—

The test: a two-proportion two-tailed z-test. The p-value is the chance of seeing a gap this large if A and B were truly equal; below 0.05 is the usual bar for “real.” The interval is the 95% range for the true difference in rates. It assumes a fixed sample decided in advance — peeking at a live test and stopping when it looks good inflates false positives.

How many do I need?

%
% relative

A 10% relative lift on an 8% baseline means catching B at 8.8%.

Visitors needed per variant—

—

Standard two-proportion power calculation at 95% significance. Halve nothing — that count isper variant, so a two-arm test needs roughly double. Smaller effects cost dramatically more traffic; that trade is the whole planning conversation.

How many holdouts?

For a rollout on people — a sales training, a new call script, a coaching program — the test is the same rep measured before and after, with some reps held back so you can see what would have happened anyway. This sizes that holdout.

trained + holdout
% of the pool

Mean and standard deviation across reps for one window of history, in whatever unit the KPI uses. SD ÷ mean is the noise you are fighting.

ρ

How much a rep’s number in one window predicts the next: CORREL over two past windows. 0.3 is churny, 0.7 is stable.

% relative

—

Holdout reps—
Trained reps—
Min. detectable lift—
Info kept vs 50/50—

—

HoldoutReps heldCatches

—

The design: randomize who waits, measure everyone the same way before and after, then compare the trained reps’ post numbers to the holdout’s after adjusting for where each rep started (ANCOVA, the sharper cousin of a plain before/after difference). A holdout of share h keeps 4·h·(1−h) of the information a 50/50 split would: 20% keeps 64%, 10% keeps 36%. Shrinking the holdout is cheap until it isn’t. The holdout gets the training after the read — nobody is denied it, only sequenced.

Did the training work?

Trained

Holdout

—

Trained change—
Holdout change—
Training effect—
p-value—

—

—

The test: each rep’s change (post − pre) is the unit. The training effect is the trained group’s average change minus the holdout’s — a difference-in-differences — checked with a Welch t-test. Whatever moved both groups (seasonality, a comp-plan change, a ramping cohort) cancels out; that is the whole reason for the holdout. “SD of per-rep change” is one STDEV over the change column. If you have the rep-level rows, regress post on pre and a trained flag instead: same answer, tighter interval.

Measurement before celebration. Knowing whether a number is real — and how much traffic it takes to find out — is the same discipline I bring to an AI feature in production.

Email meEval set sizer