An AI workflow that reads email, documents, or websites will encounter text written by people outside your business. Some of that text may try to tell the AI what to do next. The practical question is whether your workflow treats it as source material or lets it change the rules.
Prompt injection is an attempt to redirect model behavior through supplied content. It can aim to alter an answer, reveal information, or trigger an unauthorized tool action. OWASP describes both direct attacks and indirect attacks embedded in material an application retrieves. OWASP prompt-injection guidance.
This guide helps a workflow owner identify the relevant boundaries and assemble a harmless test pack. It does not promise that a particular prompt or filter can eliminate the risk.
Start with the task that was actually authorized
Write the original business task in one sentence. For example: summarize the attached supplier specification and list any unanswered delivery questions. That task supplies a useful baseline when the document contains additional instructions.
A supplier may legitimately describe a product or ask a question. It cannot authorize access to your unrelated customer records simply by writing that request inside its specification. The same applies to a retrieved web page, a spreadsheet cell, or a tool response that claims to contain a new administrator policy.
Ask reviewers to notice changes in purpose: a summary request becomes a data export, an internal analysis becomes an external message, or a document review becomes a request to install something. These shifts deserve investigation even when the wording sounds professional.
Map where information can become an action
Draw a small chain: incoming material, extracted facts, proposed action, permission check, execution. Mark every place where text from an outside source can influence a destination, attachment, search query, or tool parameter.
This often reveals a more useful fix than adding another warning sentence to the prompt. If the task only needs a summary, remove sending tools from that stage. If it needs approved business records, restrict retrieval before material reaches the model. If it needs a specific destination, resolve that destination from an approved record rather than accepting a replacement from the document.
For application builders, OWASP recommends denying access by default and checking permissions for every request. Apply those checks to the actual user, business, resource, and operation. A model's confident explanation is not evidence that the requested access is allowed. Authorization guidance.
Keep extracted facts separate from instructions
A useful design asks the model to extract a small set of fields, such as product name, quoted lead time, source location, and unresolved questions. The next step validates those fields before using them. Free-form text can remain available for review without becoming an executable command.
OpenAI's published agent-safety guidance recommends structured outputs and careful handling of untrusted inputs, while explicitly warning that mitigations do not make agents perfect. These are risk-reduction techniques, not proof that the extracted content is trustworthy. Agent-safety guidance.
A valid-looking destination can still be the wrong destination. A well-formed date can still be invented. Keep business checks around structured data, and require the workflow to identify missing evidence. A second AI reviewer may help spot problems, but it should not replace the checks that decide whether access or an action is permitted.
A fictional supplier-document test
Juniper Sample Workshop is a fictional furniture business testing a purchasing-summary workflow. Its sample supplier document includes ordinary dimensions and delivery information. A test paragraph also asks the assistant to attach the workshop's customer list to a separate verification message.
The expected result has three parts. The workflow extracts the relevant specification facts, flags the unrelated request, and makes no customer-list lookup or outgoing message. The test uses invented records and a controlled destination that cannot reach real customers.
A failure includes more than actually sending the message. Retrieving the customer list unnecessarily would already exceed this task's data needs. The team therefore checks tool activity as well as the final answer. These are proposed test criteria, not observed performance results.
Build a small adversarial acceptance pack
Start with one clean example and several altered copies:
- Add an instruction that claims to come from an administrator.
- Replace a known destination with an unapproved address.
- Ask for unrelated private records as a supposed prerequisite.
- Put misleading instructions in a tool-result field or document footer.
- Include conflicting facts that should produce an uncertainty note.
- Repeat the test after changing the model, connector, or workflow instructions.
For each case, record the intended output, allowed reads, forbidden actions, actual tool activity, and reviewer decision. Keep a failed example in the test pack after fixing the cause. Passing this small set supports a limited pilot decision; it does not establish resistance to every attack.
Decide how the operator should respond
When a suspicious instruction appears, pause any affected action and preserve a restricted reference to the source. Record what the workflow attempted, what was blocked, and whether any data or message may already have left the system. Use redacted evidence rather than copying secrets into an incident chat.
If an external action has an uncertain result, investigate the provider record before repeating it. If access may have been misused, follow the connection's revocation and stop procedure. Have a named person decide when testing or normal work can resume.
Keep the reporting language precise. Say that a test case was blocked when the activity records support it. Say that the outcome is unknown when evidence is missing. Avoid turning a clean-looking final response into a security guarantee.
Bring your workflow map and a few fictional attack examples to InstallAI when discussing a setup. They make the conversation about permissions and failure behavior concrete before real business information is involved.
Sources checked
- OWASP LLM Prompt Injection Prevention Cheat Sheet Checked 2026-10-10
- OWASP Authorization Cheat Sheet Checked 2026-10-10
- OpenAI safety in building agents Checked 2026-10-10