Key takeaway

A collection of unusual cases can test capability, but it cannot establish how often those cases occur. Report selection and denominators separately.

Choose which question the sample can answer

An evaluation may ask whether a model recognizes a specific failure pattern, or how it performs on the next month of ordinary work. A deliberately enriched failure set can help with the first question while distorting the second. Identify the claim before selecting cases, and label a selected challenge set plainly.

Scikit-learn defines accuracy as the fraction of correct predictions and recall using the positive cases actually present. Those denominators explain why a high overall score can coexist with failure to recognize the rare event. Its guidance on leakage also warns against using unavailable information in evaluation. Neither source makes an archive representative merely because its rows are scored.

Decide what counts as a failure, which unit is counted and which records have enough observation to judge. Unknown outcomes are not negative cases. Multiple notes about one breakdown are not automatically independent breakdowns.

A hypothetical denominator ledger

The invented archive below contains 1,000 eligible episodes with sufficient outcome observation: 10 failures and 990 non-failures. A separate set of 80 unresolved episodes is excluded from that eligible denominator and disclosed. The challenge sample deliberately includes all 10 failures and only 40 non-failures.

The resulting challenge-set failure share is 10 out of 50, or 20%. The eligible-archive share is 10 out of 1,000, or 1%. Reporting the former as the archive failure rate would be wrong. Neither figure estimates a future company’s rate, and the unresolved 80 cases prevent a claim about the entire original archive.

SetFailures / non-failures / unknownWhat it can support
Eligible archive10 / 990 / 0Describes this defined, observed hypothetical set
Unresolved episodesUnknown outcomes for 80 episodesCoverage limitation; not added to negative class
Selected challenge sample10 / 40 / 0Tests these selected cases; 20% is not population prevalence
Later-period evaluationNot yet inspectedCannot claim future performance or prevalence

Make missed failures visible

In the hypothetical eligible archive, a rule that predicts no failure for every episode is correct on 990 of 1,000 episodes: 99% accuracy. It detects none of the 10 failures, so failure recall is 0%. There are no predicted positives, so precision has a zero denominator and should not be reported as an informative successful detection score.

A second illustrative rule identifies six failures, misses four, and incorrectly flags 24 non-failures. Its failure recall is 6 divided by 10, or 60%; its precision is 6 divided by 30, or 20%. That tradeoff may or may not fit the operator’s review capacity. The metric alone cannot decide whether 24 extra investigations are acceptable.

Report the confusion counts, the label definition and the operating threshold alongside percentages. Ten observed failures are a small basis for judging performance; avoid false precision or guarantees about the next ten. If the buyer needs a decision under uncertainty, identify the additional observations required rather than decorating the result with more decimal places.

Prevent the curated cases from contaminating the test

Freeze the test selection rule before tuning. If reviewers inspect failures to improve labels or prompts, disclose that work and hold out a separate evaluation set when possible. Otherwise the result may describe performance on cases used during development, not an independent test. Keep all events from one episode together when the chosen claim requires episode independence.

Write a selection log with eligibility, exclusion reason, sampling method and whether cases were viewed during development. In a hypothetical log, a dramatic breakdown was selected because it was memorable, not randomly drawn. It remains a useful qualitative case, but should not be given the weight of a representative observation.

Use separate reports for the challenge set and the ordinary-work set. Mixing their scores into one headline conceals the tradeoff the reader needs to assess. If no representative set is available, say that directly and restrict conclusions to the selected cases.

Decide whether a larger evaluation is justified

Proceed when the selection is reproducible, unknown outcomes remain visible and the observed errors address the actual operational decision. Ask for a new sample when rare-event coverage is insufficient or when all informative cases were used during tuning. Pause if labels depend on unavailable follow-up or if apparent sample size mostly counts duplicated episodes.

Use the due-diligence tool to record the missing denominator or independence question. A carefully selected package can be useful without being representative; that limitation should travel with its description and results. No metric here establishes a license value, expected earnings or permission to share source records.

Tools for this decision

Diligence question builder →