Key takeaway
Define the evidence required for each outcome before counting successful cases. An administrative status and a verified result answer different questions.
Start with the question the label must answer
If an evaluator asks whether a repair prevented another failure, a record marked complete is an input to the investigation, not the answer. Write the unit first: one repair attempt on one asset, followed for a specified period. A second repair can belong to the same episode while remaining a separate attempt. Decide which of those units the label describes before anyone starts marking rows.
Zendesk’s audit reference distinguishes a field-change event from a comment event. That distinction is useful beyond tickets: a status transition records a workflow action, while an explanation or follow-up supplies different evidence. It does not establish that the underlying issue was resolved. W3C PROV provides a way to associate information with the activity and agent that produced it.
The taxonomy below is our proposed review method. It is not a vendor’s required status model. Keep the source status alongside the derived label so a reviewer can challenge the interpretation without changing the historical record.
A completed dictionary for one repair question
This hypothetical dictionary concerns recurrence of the same symptom within 30 days of an attempted repair. The window, symptom matching rule and evidence requirement are assumptions chosen for this evaluation. Another use may need a different window or outcome. Do not quietly carry these labels into a different prediction task.
A job can be completed while its outcome remains unknown. Likewise, a customer declining a proposed repair is not evidence that an attempted repair failed. Separating these states prevents an evaluator from forcing every row into success or failure just to make a binary dataset.
| Label | Required evidence | Illustrative record |
|---|---|---|
| Completed; result unknown | Work marked finished; no sufficient follow-up | Attempt A: component replaced; asset observation ends next day |
| Verified non-recurrence in window | Observation covers the defined period and symptom test | Attempt B: follow-up checks recorded through day 30; matching symptom absent |
| Observed recurrence | Linked evidence of the defined symptom returning in the window | Attempt C: same symptom documented on day 9 |
| Not attempted | Proposed action rejected or cancelled before execution | Attempt D: replacement declined; no repair performed |
| Conflicting evidence | Two relevant observations disagree; adjudication pending | Attempt E: completion note says passed; same-day test log says failed |
Keep uncertainty visible in the package
Store a label reason, the observation window actually available, and the identifiers of supporting events. Do not equate no linked follow-up with no repeat failure. The asset may have moved to another provider, the customer may have stopped reporting, or the export may omit later visits. These are possible explanations to investigate, not facts to assume about a real archive.
For an illustrative set of 100 attempts, suppose 35 have verified non-recurrence within the defined window, 15 have an observed recurrence, and 50 have an unknown outcome. These groups are mutually exclusive. Report all three counts. Calling all 85 records without an observed recurrence successful would erase the 50 unknown cases. The decision maker needs the coverage limitation even if it makes the package look less complete.
If labels were assigned retrospectively, preserve the assignment date separately from the repair and observation dates. An annotation added months later should not masquerade as a fact available to an earlier operational decision.
Resolve disagreements before expanding annotation
Start with a small, deliberately varied set that includes missing follow-up, multiple attempts, cancellations and contradictory notes. Have two reviewers assign labels independently using the same written definitions. Record disagreements by reason: different episode boundaries, different symptom matching, insufficient observation, or a genuine source contradiction. A single agreement percentage hides which rule needs repair.
In a hypothetical review, both reviewers agree on 16 of 20 cases. Three disagreements concern how to link a repeat visit; one concerns whether a sensor test covered the symptom. Clarify the linkage rule and evidence requirement, then re-review those cases. Do not simply let the senior reviewer overwrite the junior reviewer and call the original definitions reliable.
Version the dictionary and retain the adjudication rationale. If the rule changes, identify the affected labels and reprocess them explicitly. Use our due-diligence tool to turn unresolved interpretation questions into an evaluation agenda before requesting a larger sample.
Decide whether the archive supports this use
Proceed when the proposed label can be traced to permissible evidence, unknowns remain explicit, and a reader can reproduce the interpretation on the bounded sample. Narrow the use when the archive supports completion but not later success. Pause when disagreements depend on missing records that cannot be recovered within the approved scope.
This review establishes a more intelligible dataset description. It does not demonstrate buyer demand, improve a model by a stated amount, or create authority to share records. A metadata-only discussion can explain the dictionary and its coverage limits without distributing underlying narratives. For VOID, a named introduction, a sample and a license remain separate permission decisions.