Skip to content

AI security · Track C

Your AI feature is an attack surface with a natural-language front door.

Most AI security testing on the market is a scanner running a list of known jailbreak strings. That finds what everyone already knows. The failures that matter are authorisation flaws wearing a new coat — a model with broader data access than the person querying it, an agent that can be talked into using a tool it holds credentials for.

Those are access-control bugs, and access control is what this practice has spent over a decade breaking: at a top-ten US bank, against AWS itself, and on Amazon's consumer devices. The last five of those years were spent finding exactly these classes at Amazon, on shipping hardware and in vulnerability research.

AI systems

LLM and agentic application red teaming.

The OWASP Top 10 for LLM Applications and the NIST AI Risk Management Framework now sit in most enterprise procurement checklists. Testing against them is increasingly what a customer means when they ask whether your AI feature is safe.

Prompt injection

Direct and indirect. Injection through retrieved documents, tool output, and user-supplied content that reaches the model without passing a person first.

Agent and tool abuse

What an agent can be persuaded to do with the credentials, tools and permissions it holds — and what it reaches after one successful instruction.

Data leakage

System prompt extraction, training and retrieval data disclosure, and cross-tenant leakage in multi-customer deployments.

Guardrail bypass

Jailbreaks, encoding and multi-turn manipulation against whatever filtering sits in front of and behind the model.

Authorisation at the model boundary

The most common failure: a model given broader data access than the user querying it. Classic access-control testing applied to a new surface.

Supply chain

Model provenance, third-party tool and plugin trust, and what the application inherits from its dependencies.

Drivers

Who is asking you for this.

AI red teaming is moving the same way pentesting did — from a good idea to a thing somebody makes you produce evidence of. Three forces are doing the pushing, and only one of them is a law.

  1. 01

    Your enterprise customer

    The fastest driver by a wide margin, and it is contractual rather than legal. Security questionnaires now ask whether AI features have been adversarially tested and by whom. This is the same mechanism that made SOC 2 pentests effectively mandatory: no regulation required them, buyers did.

  2. 02

    NIST AI RMF — Generative AI Profile

    Voluntary, but the de facto US standard, and federal agencies and enterprise procurement increasingly expect it. Its July 2024 Generative AI Profile names prompt injection, data leakage and misuse at scale directly. The way you evidence MEASURE and MANAGE outcomes is a documented red team backlog: runs, failures, mitigations and regression results.

  3. 03

    ISO/IEC 42001 and the EU AI Act

    ISO 42001 is the AI management system standard and is moving from differentiator to table stakes in procurement. The EU AI Act is the binding one — general-purpose AI model obligations applied from August 2025, with Commission enforcement beginning 2 August 2026 and models already on the market due by August 2027. It reaches you if your product reaches the EU.

Engagements

Scope and price.

AI and LLM red team

from $16,000

Prompt injection, jailbreaks, tool and agent abuse, training and retrieval data leakage. Mapped to the OWASP Top 10 for LLM Applications and the NIST AI RMF Generative AI Profile.

  • Model and application layer
  • Agentic tool abuse
  • Data leakage
  • OWASP LLM / NIST AI RMF mapping

Typical duration · 5 days

AI red team evidence package

from $21,000

The red team engagement plus the documented evidence trail an ISO 42001 audit, a NIST AI RMF MEASURE/MANAGE mapping or an enterprise AI security review asks for.

  • Full red team engagement
  • Run, failure and mitigation log
  • Framework control mapping
  • Attestation letter

Typical duration · 7 days

Device work in particular varies a great deal with the product. The figure above assumes one device, one companion app and one backend; a fleet, several hardware revisions or a certification target will scope differently.

Scope a test