Tools · Eval harness

Eval Set Sizer

“It works on the examples I tried” is not a measurement. To claim an accuracy with a straight face, you need enough labelled cases to put a confidence interval around it. This sizes the golden set for you.

%

Your best guess at how often the feature is right. Use 50 if you truly don’t know — it’s the most demanding.

points

How tight the answer must be. ±5 means a measured 90% really sits in 85–95%.

How often the interval should contain the true accuracy.

Golden examples needed—

—

The math: normal-approximation sample size, n = z² · p(1 − p) ÷ e², rounded up.p is expected accuracy, e the margin of error, z the confidence multiplier. It assumes independent, representative examples — so sample real traffic, don’t cherry-pick. For tiny or lopsided sets the normal approximation loosens; treat the number as the floor.

Want your eval plan sanity-checked?

Leave an email and I'll send back your sample size with the three things that most often invalidate it — sampling bias, label drift, and the failure modes a single accuracy number hides. Written by me, not a sequence. Usually within a day.

Sizing the golden set is step one; building it is the work. Every AI feature I ship gets a labelled golden set and a harness that runs on each change, and the monthly bench keeps it green after the handoff. The 90-day build and the bench are the two ways to hire that work.

Work with meSee an eval in production