Key takeaway

A strong pooled result can conceal a weak facility. State which sites were unseen and whose workload the score represents.

Choose the deployment question first

A buyer may want a model to work at an entirely new facility. A test on more records from familiar facilities answers a narrower question. scikit-learn describes grouped validation in which the validation groups are absent from the paired training fold; GroupKFold implements non-overlapping groups supplied by the caller. A facility identifier can serve as that group when the deployment question concerns new facilities.

That choice does not establish that the sites represent every future customer. Record operating differences, available source systems and relevant exclusions. A facility can differ in workflow as well as record format. A group-aware split helps test a particular generalization question; it cannot repair a collection that omits the intended deployment conditions.

Do not let the larger facility hide the smaller one

This hypothetical test has 400 independently scored cases at Site A and 100 at Site B. Site A has 360 successes; Site B has 60. All figures are invented and both sites belong to the test population, not to a real seller. Pooled and equal-site summaries answer different questions.

SummaryCalculationIllustrative result
Site A success360 / 40090%
Site B success60 / 10060%
Pooled across cases420 / 50084%
Equal weight per site(90% + 60%) / 275%

The 84% pooled figure describes the illustrated case mix. It does not establish 84% performance at a new facility.

Keep the whole held-out site unseen

Define groups before model selection and keep every record from a held-out facility out of its paired training fold. Also inspect shared templates, copied episodes and preprocessing decisions that may cross the boundary. Fit learned preprocessing within the training side. Save the site grouping with the evaluation package so another reviewer can verify it.

If Site B is the sole unseen facility, its 60% result is the relevant observation for that held-out site; averaging it with familiar-site results would change the question. If multiple facilities are held out in turn, show every fold’s training sites, held-out sites, case count and result. Do not tune against the final held-out site and then describe it as untouched.

Report the intended weighting and its limits

Choose the summary that matches the decision before comparing candidates. A buyer concerned with workload across a known network may use case weighting; a buyer concerned with a typical site may also need equal-site weighting and the range. Both summaries should retain the actual denominators. A site with ten cases offers different evidence from one with thousands.

Explain exclusions and any facility with no eligible cases. Omitting a difficult site can improve the score while narrowing its meaning. Show the missing site explicitly rather than silently removing it from a denominator. Small-sample uncertainty and temporal leakage require their own checks; changing group boundaries does not resolve either.

Set a claim boundary the buyer can use

A useful conclusion says that a specified method was tested on named synthetic or permissioned site groups under a fixed procedure, with site-level results and known gaps. It does not say that all facilities will achieve the pooled score. Use that evidence to decide whether another bounded site evaluation is needed.

This decision differs from preventing future information leaking into a past decision. A time-respecting split can still include familiar facilities on both sides. The diligence question builder can capture the intended new-site use and the requested grouping evidence. No article result here demonstrates available inventory, buyer demand or a commercial premium.

Tools for this decision

Data inventory builder →Diligence question builder →