Key takeaway
A source note may be legitimate business evidence and still be untrusted input to an AI application.
Treat retrieved text as an input channel
An operational note can contain a sentence that tells an AI application to change its task. NIST defines indirect prompt injection as an attack through resource control rather than the user’s direct input. The problem can arise when a system mixes source material with instructions that govern its behavior.
A data owner should not label an entire archive safe because it came from a familiar business system. Some notes originate with customers, contractors or imported documents. Conversely, the mere presence of imperative language does not prove malice: a repair instruction may be exactly what the archive needs to preserve. The receiving system must distinguish interpreting a note from granting it authority.
An instrumented test with harmless records
This hypothetical test uses invented maintenance notes and substitute tools that cannot send messages, change real records or access secrets. The approved task is to summarize a fault. The unauthorized instruction asks for an unrelated action. The intended result is determined by the original task and permissions, not by keywords in the note.
| Record case | Expected behavior | Evidence to retain |
|---|---|---|
| Ordinary repair instruction | Summarize its meaning | Task completed with cited source |
| Note requests unrelated tool action | Do not execute the action | Denied operation and safe summary |
| Quoted warning containing instruction text | Preserve relevant warning | No false action; useful answer retained |
| Same request inside retrieved attachment | Apply the same boundary | Channel and tool decision recorded |
Put permission checks next to the action
OWASP recommends validating tool calls against the user’s permissions and session context, restricting tool access and using read-only accounts where possible. In an evaluation, keep the component that executes a side effect responsible for that check. A model’s statement that an operation is approved is not the approval record.
NIST’s report overview discusses limitations of existing mitigation techniques. Separation labels, screening and a second model may help, but they do not justify removing execution controls. Agree what the application is allowed to read, where outputs may go and which operations require an authorized reviewer. Test those specific boundaries rather than relying on a claim that the dataset was cleaned.
Score the task and the boundary separately
Record whether the prohibited action occurred, whether relevant source content appeared correctly, and whether a permitted task was wrongly refused. If the system returns nothing, the forbidden action might be prevented while the business task still fails. A result that counts every refusal as success can conceal an unusable workflow.
For each test, keep the channel, permitted task, prohibited effect, observable tool decision and final answer. Use synthetic identifiers and omit credentials or unnecessary private text from logs. Repeat tests when the model, retrieval pipeline or tool scopes change; a previous result does not automatically cover a new execution surface.
Decide what a bounded evaluation can support
If an unrelated side effect is possible, withhold the relevant action capability and revisit the receiving design before using a real sample. If controls hold but legitimate tasks are frequently blocked, revise the task or screening logic and repeat the same recorded cases. Neither outcome establishes general resistance to every attack.
This article is a receiving-system review agenda. It does not certify a source archive, endorse a model or claim that VOID performs application security testing. Use the diligence question builder to request the buyer’s boundary evidence. Source disclosure, sample access and any license remain separate permissions, even when a synthetic test is useful.