Key takeaway
A benchmark needs a defensible reference and a relevant alternative. Agreement with a convenient label is not enough to prove operational usefulness.
Choose a decision that the test can actually change
A prospective buyer may request operational records to evaluate whether an AI assistant retrieves a useful repair procedure. Before preparing a benchmark, specify the task, available context, allowed output and error consequence. If the test question is undefined, collecting a large set of answers can create impressive scores without a decision to support.
NIST’s AI Risk Management Framework calls for documented test sets, metrics and tools, suitable evaluation conditions and independent or domain assessment. It also asks organizations to consider viable non-AI alternatives. Those principles inform the review below; NIST sets no universal passing score for this particular task.
A hypothetical paired benchmark
A fictional internal pilot has 200 approved retrieval tasks. Two qualified reviewers establish a reference from the evidence available to the user. They resolve 180 tasks and mark 20 unresolved. A rules-based search tool and a proposed AI assistant then receive the same task information and are scored against the same resolved reference.
All counts are illustrative. The 20 unresolved cases remain in the coverage report; they do not become convenient wrong answers or disappear from the archive description.
| Measure | Simple search | AI assistant | Interpretation |
|---|---|---|---|
| Correct among 180 resolved tasks | 126/180 = 70% | 153/180 = 85% | A paired test on the declared resolved subset |
| Unresolved reference tasks | 20 | 20 | No correctness claim; review separately |
| Unsafe unsupported action suggestions | 2/200 | 8/200 | Separate consequential error count |
| Median completion time, entered illustration | 4 minutes | 3 minutes | Needs same timing rule; not a labor saving yet |
Inspect the disagreements, not only the average
Retain task-level results so reviewers can see where one approach succeeds and the other fails. The AI’s higher aggregate correctness in this example does not cancel its higher unsupported-action count. A decision can require improving that failure mode before proceeding, even when the average score looks favorable.
Ask the domain reviewers why a reference was chosen and which evidence would change it. Distinguish an acceptable alternative answer from a wrong answer. Keep adjudication rules fixed during the scored comparison, then record any necessary corrections as a new review version. Changing references to favor one system after seeing its outputs defeats the purpose of the comparison.
The two totals alone do not identify how many tasks only one approach answered correctly. Retain a paired result for each resolved task: both correct, search only, AI only or neither. A review of these groups can locate the changed behavior. Keep unresolved references in a fifth category rather than forcing them into the four scored groups.
NIST’s generative-AI profile specifically identifies erroneous benchmark labels as a source of unstable or misleading comparisons. Keep evidence for the reference answers and a correction history. A higher score against an unreliable reference can reward the wrong behavior; repairing the reference and rerunning both approaches is more informative than defending that score.
Keep the test conditions inside the conclusion
Match the task context to the proposed use: which documents were visible, which tools were available, which model or search version ran, and how outputs were judged. A result obtained with the answer-bearing document in the prompt cannot establish a system’s ability to locate that document from an archive. State the allowed information and record the actual setup.
Separate benchmark use from broader training or product rights. A sample supplied for a bounded evaluation does not automatically permit training, permanent storage or redistribution. If reviewers or service providers access the records, include them in the approved recipient and handling scope. Synthetic structure checks can come first; they cannot substitute for authentic task evidence.
Make the next funding decision explicit
The example supports a targeted follow-up on unsupported actions and unresolved references, followed by a new frozen paired test. It does not yet support deployment or a claim of verified labor savings. Set the follow-up scope, owner and stopping criterion before expanding the archive request.
Use buyer diligence to write the task and reference questions, then readiness to assign evidence gaps. Keep package acceptance separate from system performance: a conforming delivery can still be unsuitable for a particular application. The strongest evaluation brief states both what the supplied records must contain and what additional test evidence would justify using them.