Tools · Eval harness
Eval Set Sizer
“It works on the examples I tried” is not a measurement. To claim an accuracy with a straight face, you need enough labelled cases to put a confidence interval around it. This sizes the golden set for you.
—
The math: normal-approximation sample size, n = z² · p(1 − p) ÷ e², rounded up.p is expected accuracy, e the margin of error, z the confidence multiplier. It assumes independent, representative examples — so sample real traffic, don’t cherry-pick. For tiny or lopsided sets the normal approximation loosens; treat the number as the floor.
Want your eval plan sanity-checked?
Leave an email and I'll send back your sample size with the three things that most often invalidate it — sampling bias, label drift, and the failure modes a single accuracy number hides. Written by me, not a sequence. Usually within a day.
Sizing the golden set is step one; building it is the work. Every AI feature I ship gets a labelled golden set and a harness that runs on each change, and the monthly bench keeps it green after the handoff. The 90-day build and the bench are the two ways to hire that work.