Key takeaway

Nineteen successes out of twenty is 95% observed success. It is not precise evidence of a 95% underlying success rate.

Keep the count beside the percentage

An evaluation summary can look decisive when its numerator and denominator disappear. NIST’s handbook gives Wilson confidence bounds for a binomial proportion. SciPy documents multiple methods and defaults to the exact Clopper–Pearson method, so a reviewer must specify which calculation produced an interval. Calling every interval a Wilson interval would be wrong.

For this review, consider a binary outcome measured on independent cases from a stable target process. Those are assumptions to examine, not facts established by the formula. Repeated attempts on the same incident, cherry-picked easy examples or a shifting success definition can invalidate the intended interpretation.

Same percentage, different precision

The following hypothetical counts use a two-sided 95% Wilson interval without continuity correction, with z approximately 1.959964. Values are rounded to two decimal percentage points. These are invented evaluation outcomes, not a buyer benchmark or client result.

Illustrative sampleObserved success95% Wilson interval
19 successes / 20 cases95%76.39%–99.11%
190 successes / 200 cases95%91.04%–97.26%

What the interval actually says

Under the assumed binomial process, the Wilson procedure expresses sampling uncertainty about the underlying success proportion. It is an approximate confidence procedure, not a statement that 95% of future cases will succeed. Nor should this table be read as assigning a 95% probability to the particular fixed parameter lying inside its already-computed interval.

The first interval spans materially weaker performance than the displayed 95% point estimate. Increasing the sample count narrows this illustrated interval while leaving the point estimate unchanged. That arithmetic does not prove the larger sample is representative. If all 200 cases came from the same repeated episode, a wider collection of independent cases may be more informative than more rows.

Define the acceptance rule before collecting more

Suppose a buyer’s hypothetical rule requires a two-sided Wilson lower bound above 90%, alongside an agreed error-severity review. The 19-of-20 example fails that numerical rule; the 190-of-200 example passes it under the stated assumptions. This is an illustrative rule chosen for the example, not a recommended universal performance threshold. A one-sided rule would require a different calculation.

Write the case unit, target population, success criteria, exclusion policy, confidence level and interval method into the test plan. If results are examined repeatedly while collecting more cases, do not assume a fixed-sample confidence interpretation applies unchanged. Set a planned sample and stopping procedure with an appropriate statistical reviewer rather than stopping at the first favorable bound.

Report the failures and the sampling boundary

Preserve the failed cases and their severity; one dangerous failure can matter more to a decision than many low-consequence successes. Show counts by material subgroup where the target decision needs them, and disclose missing or dependent groups. Use the rare-event guide for coverage of unusual events and the cross-site guide for facilities absent from training.

The completed table helps a buyer decide whether a headline rate supports the intended claim or whether more suitable evidence is needed. It does not establish data value or certify a model. The diligence question builder can capture the sampling and interval questions before the parties separately approve any real sample or licensing use.

Tools for this decision

Diligence question builder →