Key takeaway

Acceptance should identify a reproducible delivery check and a responsible decision maker. A vague promise of useful data is not a test.

Specify what is being accepted

Separate delivery conformance from the buyer’s broader commercial objective. A package can meet its agreed schema while failing to improve a particular model. Conversely, a promising demonstration may conceal missing records or disallowed fields. The parties need to say which results determine acceptance and which remain evaluation observations.

DCAT 3’s version identifiers and notes can help pin the delivery under discussion. Scikit-learn’s leakage guidance explains why a result produced using unavailable information can be misleading. These references support two different questions: did the agreed package arrive, and was a claimed evaluation meaningful? They do not set contractual thresholds or guarantee business value.

Write the package identity, permitted use, test procedure, denominator, reviewer, response period and correction path in the commercial discussion. A test that neither side can reproduce is likely to become an argument about expectations rather than evidence.

An illustrative delivery test schedule

The following hypothetical schedule concerns a bounded evaluation package of 500 job rows and their permitted visit records. Every threshold is invented for this example and would require negotiation and legal review. None is a criterion of Handshake, micro1, VOID or another receiving program.

The denominator for a required field is the agreed 500 job rows, not just rows remaining after failures are silently dropped. An exclusion may be appropriate, but its treatment must be visible in the scope before the test is run. Keep the supplied package intact while recording findings in a separate result artifact.

Check / illustrative thresholdEvidenceDecision owner / cure
Identity: exact agreed packageFile identities against release manifestDelivery lead; supply intended package if different
Schema: every required column presentSchema comparison and type exceptionsBuyer reviewer; reject missing required column pending correction
Outcome evidence link: at least 475 of 500 jobsCount with defined valid evidence relationshipNamed reviewer; investigate failed joins before deciding
Excluded contact fields: none deliveredAgreed field and narrative inspection procedureOwner privacy lead; contain issue and review scope
Duplicate job keys: zero in the job tableDistinct-key reconciliationDelivery lead; correct or document intended unit
Prediction-input rule: no post-decision fieldsAvailability audit against defined prediction timeEvaluation lead; rerun affected evaluation separately

Work through a failed check

Suppose the hypothetical package has 500 job rows but only 460 valid outcome links. It fails the negotiated 475-row threshold: 460 divided by 500 is 92%, while the stated threshold is 95%. That arithmetic is not evidence that 92% is commercially unacceptable everywhere. It is evidence that this package does not meet this example’s agreed check.

The delivery lead finds 20 links failed because visit identifiers were converted incorrectly. Correcting those links would produce 480 valid relationships, or 96%, if the source evidence and authority are verified. The remaining 20 unresolved jobs still need a documented treatment. Do not label them successful simply to pass the threshold.

Record a new package version and the rerun result. The buyer should receive the correction explanation, not only a replacement file. If the original defect also affected a reported model score, determine which analysis needs rerunning rather than treating a delivery pass as proof that the score remains valid.

Define who decides and what happens next

Agree how findings are communicated, how disputed counts are checked and when a corrected delivery can be retested. Name an accountable reviewer on each side. A system-generated pass is evidence for that reviewer, not a substitute for a decision when scope or confidentiality is disputed.

Separate partial acceptance, a request for correction, rejection and silence. Do not assume that a recipient’s download means acceptance, or that silence means a payable amount exists. Payment triggers belong in their own terms and may occur before or after particular acceptance events. Counsel should review the effect of any proposed acceptance-by-silence clause.

Avoid open-ended obligations to fix any issue the buyer later discovers. Describe the applicable package, permitted evaluation, response process and unresolved questions. A delivery test schedule helps the discussion but does not replace a license, privacy assessment or agreed remedy.

Use the schedule before spending on preparation

Ask whether the proposed tests are measurable with the available source records and proportionate to the bounded evaluation. If the buyer requires a history the owner never collected, that is a scope mismatch to resolve before reconstruction. If a test concerns model performance, pin the method, inputs and exclusions rather than promising improvement in the abstract.

Use the due-diligence tool to identify missing evidence and the offer-comparison tool to record the actual acceptance and payment terms when an offer exists. These hypothetical checks assign no data value and promise no sale. For VOID, a permissioned introduction does not grant sample access or commit the company to meeting a receiving program’s future delivery requirements.

Tools for this decision

Offer comparison →Diligence question builder →