Why prompt injection is a security problem
A large language model (LLM) processes instructions and other content within its input context. An application may ask it to summarise a document, answer from a knowledge base or suggest an action. The document or retrieved text can itself contain instruction-like material, even though its author has no authority over the application.
The security question is whether that material gains influence beyond its permitted role. An ordinary request for a shorter summary is not automatically an attack. A document that causes the assistant to abandon its task, disclose restricted information or request an unauthorised action creates a different concern.
Do not judge impact solely by whether the answer sounds alarming. A model can invent a claim about having accessed data or performed an action. Evidence must establish what actually happened in retrieval, application decisions or connected tools. Equally, a quiet change to a recommendation may matter if the business relies on it.
Direct and indirect prompt injection
The distinction concerns where the instruction enters the system. It does not determine severity: consequences depend on the permissions and controls around the model. OWASP’s prompt-injection guidance describes both entry routes.
| Route | Entry point | Defensive question |
|---|---|---|
| Direct | Input supplied through the user-facing prompt or conversation | Can a user redirect behaviour beyond their permitted task or authority? |
| Indirect | Documents, websites, messages or tool results processed by the application | Can an external author influence behaviour as if they were an authorised instruction source? |
For direct injection, testing should distinguish legitimate customisation from a bypass of application rules. A permitted user may change tone or format; that does not give them authority to access another user’s information or change approval requirements.
Indirect injection matters because the person operating the assistant may be entirely legitimate. They may ask for a normal summary of a document containing hostile instructions. Reviewing only the user’s typed request would miss the point where external content enters the workflow. Retrieved content does not become trustworthy simply because a trusted tool delivered it.
Why a secret system prompt is not access control
A system prompt can describe the assistant’s purpose and intended behaviour. It should not hold credentials or serve as the only mechanism that decides which records a user may access. Those decisions belong in application and service controls that enforce identity and permissions independently of the model.
Disclosure of prompt wording is not automatically equivalent to a sensitive-data breach. Its significance depends on what it contains and what that disclosure enables. Conversely, hiding the prompt does not repair an application that trusts the model to approve privileged actions. OWASP’s system-prompt leakage guidance explains this distinction.
Instructions telling the model to ignore suspicious text can be useful as one layer. They are not a substitute for controls that prevent a disallowed operation even if the model requests it. Evaluate the action or data boundary, rather than treating a refusal message as proof that all paths are blocked.
How retrieval and agents extend the consequences
Retrieval-augmented generation (RAG) supplies relevant information from external knowledge sources to a model at response time. A knowledge base may include uploaded files, shared documents or material imported from other systems. An assessment should consider who can contribute or alter that content and how retrieval selects it.
Permission-aware retrieval remains necessary even when the content contains no malicious instructions. If the application retrieves a restricted record for the wrong user, an access-control defect already exists. Prompt injection may interact with that defect, but the two should not be confused or reported as the same finding.
An AI agent is a system that can use tools or take actions on behalf of a user or workflow. Tool access can turn an answer-manipulation problem into an action problem. The important questions include which operations are available, whose identity the tool uses and whether an independent control approves consequential actions.
Excessive functionality, permissions or autonomy can increase the impact. A tool able to read a document need not also be able to export, modify or delete it. OWASP’s excessive-agency guidance connects these design choices to potential harm.
Hypothetical example
A procurement assistant summarises fictional supplier documents. One test document contains instruction-like material that attempts to influence the recommendation. A second version of the workflow also allows the assistant to prepare a supplier record in a test system. These examples use synthetic content and no real supplier information.
- Read-only workflow: the concern is whether the summary or recommendation follows the document’s instructions instead of the user’s task. Evidence compares the source, intended criteria and generated result. A changed answer is an integrity issue; it does not prove that any database was changed.
- Connected workflow: the concern is whether the document influences a tool request outside the approved task. Tool and application logs should show whether an operation was requested, blocked, approved or completed. A model’s statement that it acted is insufficient evidence.
- Control to examine: the receiving service checks the user’s authority and allowed action independently, and requires meaningful confirmation before a consequential change.
- Retest: vary the synthetic documents and conversation context, confirm that unauthorised actions remain blocked, and check that legitimate summaries and approved actions still work.
The same document can have different consequences in these workflows because the available actions differ. Prioritise the business effect and the control that permits it, rather than assigning severity from the prompt alone.
Controls that reduce exposure and limit impact
There is no single perfect prompt that establishes a reliable security boundary across every model, application and input. Use several controls, and test the complete workflow rather than evaluating the model in isolation.
Restrict what the system can access and do
Give tools only the operations and permissions needed for their task. Apply user-specific authorisation to each request, rather than using a broadly privileged shared identity without further checks. Enforce document and tenant boundaries before restricted information reaches the model.
Keep external content in its intended role
Track where documents and tool results came from, control ingestion permissions and separate external material from trusted application instructions where practical. Labelling and delimiters can help interpretation, but they do not guarantee that a model will ignore embedded instructions.
Validate outputs and action parameters
Treat generated arguments and text as untrusted input to the receiving component. Check permitted operations, destinations, values and ownership. Apply context-appropriate output encoding when displaying content. An application programming interface (API) must enforce its own rules even when a request was proposed by AI.
Make approval and monitoring meaningful
For consequential actions, show the user the actual destination, data and effect being approved. A vague confirmation step may not help them detect a manipulated request. Record relevant retrieval and tool events, protect logs containing sensitive information, and establish a response when suspicious activity occurs.
Filters and model safeguards may reduce some unwanted behaviour, but can miss unexpected inputs or block legitimate work. Measure both security outcomes and usability. Existing web application and API controls remain part of the defence; AI does not replace them.
How to assess controls safely
Testing requires explicit written authorisation, a defined system boundary and rules of engagement. Agree test identities, synthetic records, permissible tool effects, request and spending limits, escalation contacts and stop conditions. Confirm permissions for third-party services before including them.
Map where untrusted material enters, what the model receives and where its output can cause an effect. Define expected behaviour for each role. A useful scenario states the boundary being tested and the evidence needed to determine success or failure, rather than merely collecting surprising responses.
Repeat meaningful scenarios and record variability. Capture relevant configuration and model versions where available, along with retrieval and tool traces. Separate a demonstrated disclosure or action from an unverified statement generated by the model. Findings should identify the failed control and who can change it.
After remediation, repeat the original scenario and related variations. Check authorised workflows as well as blocked actions, and retain appropriate regression cases for future changes. Security retesting and validation provides a defined way to check corrections without promising that every future input has been covered.
What a business should ask before deployment
- Which people or external sources can influence the content the model reads?
- What information can it retrieve, and how are the requesting user’s permissions enforced?
- Which tool actions are possible, and which require independent approval?
- Can the team distinguish a proposed action from an executed one in the logs?
- Who owns fixes when the control sits in a connector, application or third-party service?
- Which model, data-source or permission changes trigger reassessment?
These questions help turn a broad concern about prompt injection into an actionable review. For the wider assessment process, including reporting and preparation, read our guide to AI security testing for UK businesses.
Frequently asked questions
Is prompt injection the same as a normal prompt?
No. A normal request operates within the user’s permitted task. Prompt injection concerns untrusted instructions redirecting intended behaviour, potentially crossing a security or workflow boundary.
Can a system without tools still be affected?
Yes. It may generate manipulated answers or expose information available in its context. Tools introduce further possible effects, but are not a prerequisite for harmful output.
Does RAG prevent prompt injection?
No. RAG adds information sources; it does not automatically establish trust in their contents. Retrieval permissions and treatment of external text still need controls and testing.
Will changing the system prompt fix it?
It may improve some behaviour, but it should not be the only correction for a data or action-control failure. Fix the relevant application, retrieval or tool boundary and retest.
Does every successful redirection have the same severity?
No. Severity depends on the demonstrated effect, reachable data, permissions and business context. An altered answer and a completed unauthorised action require different evidence and impact analysis.
When should we retest?
Retest after relevant fixes and reconsider coverage after changes to models, retrieval sources, tools, permissions or workflows. Maintain regression cases for important previous failures.
AUTHORISED SECURITY TESTING
Understand the risk in your AI workflow.
Discuss the content, data and actions your system handles, and define a controlled assessment of the boundaries that matter.