Skip to content

Home Expertise / Secure AI

AI / AI-05

AI system security testing and validation

Assess a complete AI application: model, data, connectors, policies, tools and users. Tests examine information leakage, malicious instructions received by the system and limits on possible actions.

WHEN IT HELPS

A focused response
to a defined need.

Assistant ready for deployment, new sensitive connector, action-capable agent or customer seeking security evidence beyond a content filter.

AT A GLANCE

Family
Secure AI

Engagement
Advisory and implementation

Reference
AI-05

SCOPE & OUTCOMES

What the engagement covers.

Scope

  • Authorised abuse scenarios
  • synthetic datasets
  • permissions
  • prompt injection
  • retrieved data
  • tool actions
  • privacy
  • robustness
  • cost
  • repeatability and regression testing

Deliverables

  • Test plan
  • scenarios
  • sanitised evidence
  • reproducible results where possible
  • contextual severity
  • remediation
  • statistical limitations and retest plan

Acceptance evidence

Scenarios and versions are documented; sensitive tests avoid unnecessary real data; results distinguish confirmed defects from non-reproduced behaviour; fixes are verified against the agreed set.

DELIVERY

How the work is structured.

Approach

Choose a use case; classify data and access; design and pilot; test privacy, actions and cost; decide rollout and monitoring.

Prerequisites & responsibilities

Customer: business sponsor, data owners, identity team, DPO/legal where needed and budget. Provider: architecture and tests; customer retains approval of sensitive use.

Scope factors

Uses, users, models, data, connectors, permissions, actions, volumes and hosting. Separate project, licenses, tokens, search, storage and operations; no claimed savings without measurement.

Questions to clarify

Which data or actions are prohibited? Can testing use an isolated environment? How is success measured and testing repeated after a model change?

IMPORTANT BOUNDARIES

Behaviour can vary across runs and versions. A test score is not certification; testing must respect authorisation and exclude out-of-scope third-party systems.

Permissions, provider data use, retention, residency and cost enforcement are assessed separately. Technical features and applicable obligations are checked for the chosen offering and use case.

IN PRACTICE

Illustrative situations.

These examples describe possible engagements and target outcomes. They are not customer references or achieved results.

Scenario 01

An assistant summarises third-party documents. Project: test whether content can make it ignore limits or reveal information. Target outcome: identified defects and controls tested on representative cases, without universal protection claims.

Scenario 02

An agent can create tickets and send messages. Project: verify recipients, content and human approvals. Target outcome: prohibited actions are rejected in testing; uncontrolled behaviour leads to reduced scope.

Technology and reference context

References: OWASP GenAI and NIST GenAI Profile; system-appropriate evaluation tools, versioned test sets and synthetic data.

The final technology set is agreed during scoping, based on interoperability, licensing, access rights and operating requirements.

CONNECTED SERVICES

Build the next step.

These services can complement the engagement. They are not automatically included.

START A CONVERSATION

Make the scope clear.

We will clarify the objective, dependencies and responsibilities of this service before proposing delivery.