AI / AI-05
AI system security testing and validation
Assess a complete AI application: model, data, connectors, policies, tools and users. Tests examine information leakage, malicious instructions received by the system and limits on possible actions.

WHEN IT HELPS
A focused response
to a defined need.
Assistant ready for deployment, new sensitive connector, action-capable agent or customer seeking security evidence beyond a content filter.
SCOPE & OUTCOMES
What the engagement covers.
Scope
- Authorised abuse scenarios
- synthetic datasets
- permissions
- prompt injection
- retrieved data
- tool actions
- privacy
- robustness
- cost
- repeatability and regression testing
Deliverables
- Test plan
- scenarios
- sanitised evidence
- reproducible results where possible
- contextual severity
- remediation
- statistical limitations and retest plan
Acceptance evidence
Scenarios and versions are documented; sensitive tests avoid unnecessary real data; results distinguish confirmed defects from non-reproduced behaviour; fixes are verified against the agreed set.
DELIVERY
How the work is structured.
Approach
Choose a use case; classify data and access; design and pilot; test privacy, actions and cost; decide rollout and monitoring.
Prerequisites & responsibilities
Customer: business sponsor, data owners, identity team, DPO/legal where needed and budget. Provider: architecture and tests; customer retains approval of sensitive use.
Scope factors
Uses, users, models, data, connectors, permissions, actions, volumes and hosting. Separate project, licenses, tokens, search, storage and operations; no claimed savings without measurement.
Questions to clarify
Which data or actions are prohibited? Can testing use an isolated environment? How is success measured and testing repeated after a model change?
IMPORTANT BOUNDARIES
Behaviour can vary across runs and versions. A test score is not certification; testing must respect authorisation and exclude out-of-scope third-party systems.
Permissions, provider data use, retention, residency and cost enforcement are assessed separately. Technical features and applicable obligations are checked for the chosen offering and use case.
IN PRACTICE
Illustrative situations.
These examples describe possible engagements and target outcomes. They are not customer references or achieved results.
Scenario 01
An assistant summarises third-party documents. Project: test whether content can make it ignore limits or reveal information. Target outcome: identified defects and controls tested on representative cases, without universal protection claims.
Scenario 02
An agent can create tickets and send messages. Project: verify recipients, content and human approvals. Target outcome: prohibited actions are rejected in testing; uncontrolled behaviour leads to reduced scope.
Technology and reference context
References: OWASP GenAI and NIST GenAI Profile; system-appropriate evaluation tools, versioned test sets and synthetic data.
The final technology set is agreed during scoping, based on interoperability, licensing, access rights and operating requirements.
CONNECTED SERVICES
Build the next step.
These services can complement the engagement. They are not automatically included.
START A CONVERSATION
Make the scope clear.
We will clarify the objective, dependencies and responsibilities of this service before proposing delivery.
