What is AI security testing?
An AI security assessment asks whether a system can be made to behave outside its intended permissions and business rules. For a language-model application, that includes the instructions it follows, the information it retrieves, the outputs it produces and any actions it can request. A convincing answer is not necessarily an authorised or safe answer.
This guide focuses on language-model applications, including retrieval-augmented generation (RAG), which supplies a model with relevant material from external knowledge sources at answer time. It also covers AI agents: systems that can use tools or take actions on behalf of a user or workflow. Other machine-learning systems may require different testing methods.
AI security testing for UK businesses should start with the deployment, the information it can access and the consequences of its actions.
Why traditional penetration testing is not always enough
Conventional application and network testing remains essential. Weak authentication, broken access control, exposed services and insecure APIs do not disappear when a product adds AI. A secure model cannot compensate for an application that lets one customer read another customer’s records.
Untrusted text, retrieved material and model-selected tool calls introduce additional attack paths. These behaviours can require repeated, context-aware checks within an explicitly defined penetration testing engagement or a focused AI assessment.
What does AI security testing examine?
Prioritise the following areas according to data sensitivity, user roles, external inputs and permitted actions; not every system needs every test.
Prompt injection and indirect prompt injection
Prompt injection occurs when untrusted input is interpreted as instructions in a way that redirects or overrides the application’s intended behaviour. Ordinary user-directed changes, such as asking for a shorter answer within permitted rules, are not automatically attacks.
Direct prompt injection arrives through user input. Indirect prompt injection arrives through material the application processes, such as documents, websites or tool results. Tests examine whether those instructions cross an intended boundary, not simply whether a response is unusual. OWASP’s prompt-injection guidance explains the distinction.
Sensitive data and system prompt exposure
Testing can check whether responses, conversation history or retrieval results expose information to the wrong user. Use synthetic records to examine cross-user and cross-tenant boundaries. A leaked system prompt is not automatically equivalent to a data breach; significance depends on what it reveals and enables. Secrets and authorisation decisions should not rely on keeping prompt text hidden. See OWASP guidance on sensitive information disclosure and system prompt leakage.
Retrieval integrity and RAG security
Assess whether a user can retrieve restricted material, influence the knowledge base or cause untrusted retrieved text to be treated as an instruction. Retrieval manipulation and poisoning risks concern the content and retrieval pipeline; they do not necessarily mean that the underlying model weights have changed.
Agent permissions and connected tools
An agent’s effective permissions matter more than a promise in its prompt. Tests should examine which actions tools expose, whose identity they use and whether consequential operations require independent approval. A model that can draft a message should not automatically be able to send it. OWASP’s excessive-agency guidance connects risk to functionality, permissions and autonomy.
Application boundaries and integrations
Authentication, authorisation, sessions and API controls around AI features still need testing. Check the boundary between model suggestions and trusted application decisions, including business rules such as who may approve a request. The underlying web application and API security assessment should cover these controls where in scope.
Output handling and business logic
Generated text, code or structured arguments must not become trusted commands merely because a model produced them. Assess how downstream systems validate and use outputs, and whether legitimate features can be combined into an unauthorised workflow. OWASP’s output-handling guidance describes why the receiving application’s controls matter.
What is AI red teaming?
AI red teaming is a structured adversarial exercise that explores how an AI system might fail under deliberate challenge. In a cybersecurity engagement, the objectives might include crossing a data boundary, manipulating a tool decision or bypassing an approval step. It is broader than running a fixed prompt checklist.
Red teaming can reveal behaviour that normal functional testing misses, but its results are bounded by the scenarios and access tested. Wider AI red-teaming programmes may also assess safety or other harms; those activities should not be assumed to be included in a security-only scope. The NIST Generative AI Profile discusses adversarial evaluation within a wider risk-management approach.
How RAG and AI agents change the attack surface
For RAG, follow the data from ingestion through permission-aware retrieval to the model’s response. Permissions must be enforced for the requesting user before restricted material reaches the model. Review how document permissions are retained when content is indexed and changed. OWASP’s retrieval and embedding guidance covers these access and integrity risks.
For agents, trace the next step: which tool action can follow a response, under whose identity, and with what approval? A restriction at the chat interface is insufficient if a downstream tool can still act with broader permissions.
Hypothetical example
A document assistant returns a synthetic Finance record to an Operations test user. The example uses invented departments and test data only.
- Intended boundary: the Operations user may retrieve Operations documents, but no Finance-only records.
- Failed control: the retrieval service uses a shared identity and omits the requesting user’s document-permission filter.
- Evidence: record the test user’s role, the Finance record’s access rules, a unique synthetic marker, the retrieval trace and the returned answer. The trace must show that the restricted record was retrieved; a plausible invented answer alone would not prove disclosure.
- Remediation direction: enforce permission-aware retrieval, retain document access rules during indexing and deny access when permission information is missing.
- Retest: repeat the request and variants across permitted and restricted roles, including after document permissions change. Confirm that restricted content is excluded while legitimate access still works.
What should a business prepare before AI security testing?
- Architecture and data flows: show models, retrieval sources, interfaces and connected systems.
- User and test roles: provide representative accounts and their intended access boundaries.
- Authorised tools and actions: identify permitted operations, limits and approval requirements.
- Representative test environment: include realistic configuration and synthetic data where practical.
- Third-party permissions: confirm any required approvals for providers and connected services.
- Application and security logs: arrange appropriate access to retrieval, tool-call, identity and audit records.
Testing safeguards control how the assessment is conducted, including action limits, request volumes and stop conditions. Availability, cost consumption and resistance to abuse are separate security concerns: for example, whether repeated requests or agent loops can exhaust capacity or incur excessive charges. Include those risks in the assessment where relevant, within agreed safe limits.
What happens during an AI security assessment?
Testing begins only with explicit written authorisation, an agreed scope and rules of engagement. A practical engagement usually follows these stages:
- Scope. Identify the systems and owners, permitted activities, exclusions, test windows, data handling and stop conditions. Confirm any third-party testing permissions.
- Map components and integrations. Record models, interfaces, data sources, retrieval services, tools, identities and external dependencies. Establish what each component can access.
- Threat model. Agree realistic misuse scenarios and the assets or business decisions that need protection. Prioritise meaningful paths rather than isolated novelty prompts.
- Test safely. Use test accounts, synthetic information and controlled environments where practical. Agree limits on actions, request volume and costs; keep destructive activities outside scope unless specifically authorised.
- Validate findings. Reproduce important behaviour, document prerequisites and distinguish a demonstrated security failure from an unsupported possibility. Record variability when results are not deterministic.
- Report. Explain the affected boundary, evidence, potential impact, severity and practical changes. State coverage and limitations clearly.
- Remediate and retest. Implement controls and repeat relevant tests under agreed conditions. Check that a fix addresses the underlying weakness rather than only one phrasing.
What should an AI security testing report include?
A useful report lets a technical team understand the finding and a business owner decide what to do next. Each material finding should identify the affected component, prerequisites, observed behaviour, supporting evidence, severity and potential business impact. Reproduction information should include relevant roles, configuration and model or application versions where available.
Remediation should explain which control needs to change and who can act on it. Sensitive prompts, outputs and logs need appropriate handling. The report should separate confirmed findings from limitations, and retest records should show what was checked and the outcome. Security retesting and validation can provide evidence for closure without implying that every possible attack has been eliminated.
Does my business need AI security testing?
Consider an assessment when AI processes sensitive information, serves untrusted users or can influence consequential actions. Typical candidates include customer-facing chatbots, internal copilots, RAG applications and agents connected to business tools or APIs. Even an internal assistant may cross departmental or customer access boundaries.
A useful first step is to ask: what information can this system reach, whose permissions does it use, and what happens if its answer or action is wrong? A low-impact prototype using public data may need a lighter review than an operational agent with write access. Ask a provider to explain coverage, required access, safe testing arrangements and deliverables before selecting an engagement.
AI security testing vs penetration testing
Traditional penetration testing can include AI-specific attack paths when these are explicitly in scope. The comparison below describes typical emphasis, not mutually exclusive disciplines. Choose coverage around the system’s risks; neither service label alone guarantees that every relevant layer is tested.
| Consideration | AI security testing | Penetration testing |
|---|---|---|
| Main question | Can model behaviour and connected workflows cross intended boundaries? | Can weaknesses in the agreed application, network or infrastructure scope be exploited? |
| Typical areas | Prompt handling, retrieval, data access, agent tools and approval controls | Authentication, access control, configuration, software weaknesses and attack paths |
| Evidence | Prompts, outputs, tool activity and demonstrated boundary failures | Reproducible technical findings and controlled evidence of impact |
| Shared requirements | Explicit written authorisation, agreed scope, practical remediation and retesting | Explicit written authorisation, agreed scope, practical remediation and retesting |
How often should AI systems be tested?
Use risk and change to set the cadence. Assess significant systems before launch, then reconsider coverage after model changes, new tools or data sources, altered permissions, material architecture changes or a security incident. Higher-risk deployments may justify periodic reassessment alongside ongoing monitoring.
There is no single testing interval that suits every AI system. A provider’s model update can change behaviour even when your own code stays the same. Record the version and configuration tested, maintain regression scenarios for important findings and decide who must trigger reassessment. A passing assessment is a point-in-time result, not a permanent assurance or a universal legal compliance certificate.
Frequently asked questions
Is AI security testing the same as penetration testing?
Not exactly. AI security testing focuses on AI-enabled behaviour and integrations, while penetration testing can include these attack paths when explicitly scoped. Ask what will be tested rather than relying on the engagement’s name alone.
Can prompt injection be completely prevented?
No blanket guarantee is credible for every deployment. Reduce exposure with limited permissions, trusted application-side controls, careful data boundaries and approval for consequential actions. Test those controls rather than relying only on instructions telling the model to behave safely.
Can AI security testing damage a live system?
Poorly controlled testing can affect data, availability or costs. Agree explicit written authorisation, scope, test accounts, action limits and stop conditions. Prefer a representative test environment where practical; any production testing needs specific safeguards.
Do you need access to the AI model itself?
Not always. Many application-level assessments can test a deployed interface without model weights. Architecture information, test roles and logs can improve coverage. Model-internal evaluation requires a different scope and access arrangement.
Can RAG applications be security tested?
Yes. An agreed assessment can examine retrieval permissions, knowledge-source integrity, tenant separation and the handling of untrusted retrieved content. Test coverage should reflect the application’s actual retrieval pipeline.
Can AI agents be tested?
Yes. AI agent security testing can examine tool permissions, action approval, identity propagation and behaviour across multi-step workflows. Use controlled targets and safeguards appropriate to the actions the agent can perform.
What happens after vulnerabilities are found?
Findings should be prioritised, assigned to an owner and remediated. Agreed retesting then checks the fixes and records remaining limitations. Keep useful regression cases so later changes can be checked against previous failures.
AUTHORISED SECURITY TESTING
Turn questions into an agreed test plan.
Explore our AI Security Testing service to discuss your application, its data and connected tools, and agree an appropriate assessment.
For wider coverage, explore our cybersecurity testing services.