Key takeaway
An export timestamp cannot tell you what was known when a decision was made. Reconstruct availability before trusting an evaluation score.
Write a prediction-time contract
State the decision as a sentence with a time: at the first dispatch, predict whether another visit will be needed within 14 days. Then list the information the operator could legitimately have accessed at that point. A current export may contain the entire completed case, including outcomes and notes added later. Those later fields may be useful labels while being invalid inputs to that earlier prediction.
Scikit-learn describes leakage as using information unavailable at prediction time and warns that it can inflate evaluation results. Its guidance also separates fitting a transformation on training data from applying that fitted transformation to test data. These are related checks: later case information and test-informed preparation can each make a result look better than the intended deployment permits.
Record the target, decision time, unit and observation horizon in the evaluation brief. If reviewers cannot agree on when the prediction is made, they cannot agree on which inputs are eligible. A vague goal such as improve service efficiency is not a test specification.
An illustrative availability timeline
This hypothetical service case asks for a prediction at 09:00 on Monday. Event occurrence, system entry and export time are different columns. A note describing a Monday inspection but entered on Thursday was unavailable on Monday unless another contemporaneous record demonstrates otherwise. The table makes the admissibility decision explicit.
Keep excluded fields in a separately controlled outcome or audit layer when appropriate. Do not delete evidence from the owner’s original records to make a cleaner training file. If historical availability cannot be reconstructed, narrow the evaluation claim or exclude that field; an invented timestamp does not solve the problem.
| Field | First available | Use at Monday 09:00 |
|---|---|---|
| Reported symptom | Monday 08:30 | Candidate input; verify original report |
| Dispatch category | Monday 08:45 | Candidate input; preserve original category |
| Inspection finding | Monday 13:00 | Exclude from dispatch-time inputs |
| Final repair summary | Thursday 16:00 | Outcome/audit layer only |
| Repeat visit indicator | Following week | Potential label after defined horizon |
| Exported last_modified | Month-end export | Does not establish first availability |
Separate time from case identity
A chronological split prevents some future-to-past contamination, but one episode can still appear on both sides of the boundary. If a long-running case starts in the training period and concludes in the test period, its duplicated narrative or linked visits may expose the answer. Decide whether the intended use is forecasting later work for familiar assets or generalizing to unfamiliar assets; those are different tests.
An illustrative plan trains on episodes opened before April, validates on episodes first opened in April, and tests on episodes first opened in May. A case spanning March and April is assigned as a whole, and unresolved outcome horizons are excluded from scored results until observable. This is a proposed plan, not a universal split recipe.
Scikit-learn’s TimeSeriesSplit documents ordered folds and an equal-spacing condition for comparable fold durations. Operational events often arrive irregularly, so choosing a fixed count of events does not automatically create equal calendar windows. Its gap parameter counts samples. If the task requires a 14-day buffer, implement and verify a calendar rule rather than interpreting 14 samples as 14 days.
Audit preparation as well as the split
Write down which data were used to choose vocabulary, fill missing values, select fields, deduplicate cases and tune thresholds. Fit learned preparation steps within the training partition, and preserve a held-out test set until the procedure is settled. A clean time boundary is not enough when the entire export was used to select the most predictive fields.
Create a leak register with the field, suspected route, evidence and correction. In a hypothetical audit, the field final_action contains the repair chosen after inspection. Excluding it is necessary for the dispatch-time question. The team also finds that a summary generated after case closure restates the outcome in plain language, so removing the obvious status column alone would leave the answer embedded in text.
After corrections, rerun the same defined evaluation and report what changed. Do not describe the difference as a data-value increase or decrease. It is evidence about an evaluation method under its stated assumptions.
Use the result to make a bounded decision
Approve a larger evaluation only when eligibility can be explained per field, related episodes cannot cross the chosen boundary unexpectedly, and preparation can be reproduced. If the archive only preserves the final state, it may support retrospective analysis while failing the historical prediction question. That is a useful conclusion; it prevents buying a promise the records cannot substantiate.
Bring the timeline, split rules and remaining uncertainties to the due-diligence discussion. No score on its own proves a commercial return, acceptable disclosure or the quality of another company’s records. VOID can help scope a permissioned introduction; access to samples and training rights requires separate review and agreement.