Key takeaway
Nineteen successes out of twenty is 95% observed success. It is not precise evidence of a 95% underlying success rate.
Keep the count beside the percentage
An evaluation summary can look decisive when its numerator and denominator disappear. NIST’s handbook gives Wilson confidence bounds for a binomial proportion. SciPy documents multiple methods and defaults to the exact Clopper–Pearson method, so a reviewer must specify which calculation produced an interval. Calling every interval a Wilson interval would be wrong.
For this review, consider a binary outcome measured on independent cases from a stable target process. Those are assumptions to examine, not facts established by the formula. Repeated attempts on the same incident, cherry-picked easy examples or a shifting success definition can invalidate the intended interpretation.
Same percentage, different precision
The following hypothetical counts use a two-sided 95% Wilson interval without continuity correction, with z approximately 1.959964. Values are rounded to two decimal percentage points. These are invented evaluation outcomes, not a buyer benchmark or client result.
| Illustrative sample | Observed success | 95% Wilson interval |
|---|---|---|
| 19 successes / 20 cases | 95% | 76.39%–99.11% |
| 190 successes / 200 cases | 95% | 91.04%–97.26% |
What the interval actually says
Under the assumed binomial process, the Wilson procedure expresses sampling uncertainty about the underlying success proportion. It is an approximate confidence procedure, not a statement that 95% of future cases will succeed. Nor should this table be read as assigning a 95% probability to the particular fixed parameter lying inside its already-computed interval.
The first interval spans materially weaker performance than the displayed 95% point estimate. Increasing the sample count narrows this illustrated interval while leaving the point estimate unchanged. That arithmetic does not prove the larger sample is representative. If all 200 cases came from the same repeated episode, a wider collection of independent cases may be more informative than more rows.
Define the acceptance rule before collecting more
Suppose a buyer’s hypothetical rule requires a two-sided Wilson lower bound above 90%, alongside an agreed error-severity review. The 19-of-20 example fails that numerical rule; the 190-of-200 example passes it under the stated assumptions. This is an illustrative rule chosen for the example, not a recommended universal performance threshold. A one-sided rule would require a different calculation.
Write the case unit, target population, success criteria, exclusion policy, confidence level and interval method into the test plan. If results are examined repeatedly while collecting more cases, do not assume a fixed-sample confidence interpretation applies unchanged. Set a planned sample and stopping procedure with an appropriate statistical reviewer rather than stopping at the first favorable bound.
Report the failures and the sampling boundary
Preserve the failed cases and their severity; one dangerous failure can matter more to a decision than many low-consequence successes. Show counts by material subgroup where the target decision needs them, and disclose missing or dependent groups. Use the rare-event guide for coverage of unusual events and the cross-site guide for facilities absent from training.
The completed table helps a buyer decide whether a headline rate supports the intended claim or whether more suitable evidence is needed. It does not establish data value or certify a model. The diligence question builder can capture the sampling and interval questions before the parties separately approve any real sample or licensing use.