Prompt injection testing for RAG and AI agents
What to test in an AI feature that reads customer documents or calls tools, and what to ask whoever tests it.
Treat the model as untrusted
Anyone can make a chatbot say something silly, and that’s a poor reason to pay for testing. Test something narrower. If a customer uploads a document with hidden instructions in it, can your AI feature read another customer’s files or send an email nobody approved?
Injected instructions arrive in two ways. A user can type them into the chat, or they can sit inside something the model reads, such as an uploaded file or a web page. The second kind is harder to catch because the person asking the question may have no idea the instructions are there.
We’d treat anything the model produces as a request from an untrusted user. Ordinary code, using the signed-in user’s own permissions, decides whether it goes ahead.
Test around business consequences
Testing should use made-up customer accounts and actions that can be undone. Nobody needs your real data or a real outgoing email to see whether a boundary holds.
Take a support assistant that summarises uploaded contracts. Someone hides a line in a contract telling the assistant to fetch other customers’ records and email them out. Whether that works has little to do with the model’s reply. It depends on whether retrieval only ever returns documents the user may see, and whether the email tool checks permission for itself.
Agents add a further problem. Each tool call can look fine on its own while the sequence exports data to an outside address.
Controls that limit the damage when prompts fail
Filters and system prompts make bad output less likely. Access still has to be checked in code each time a document is retrieved or a tool runs. Two failures are common. A tool runs with the application’s full credentials, so whatever the model asks for, it gets. Or an approval button says Continue without showing what will be sent or changed.
These are the questions we’d ask your team, or any provider testing the feature.
- Does retrieval filter documents by the signed-in user before the model sees them?
- Which tools can the AI call, and whose permissions do they run with?
- Which actions need a person to approve them, and does the approval screen show the exact target?
- Do the same limits apply when the agent runs on a schedule or through the API?
What good evidence looks like
Each test should record the input and what the system did with it. After a fix, the same test runs again and the logs show the request being refused.
We’d also test the paths around a fix. A limit added to the chat window may not apply to the API or a scheduled agent.
Which part of your company this is about
This note is about your AI features: what they can read and do. Our AI and agent security assessment tests these boundaries in your product.
