AI & Agentic Red Team
Offensive testing of LLMs and AI agents — prompt injection, jailbreaks, tool misuse, agent goal hijack, memory/context poisoning, model extraction, and agent identity abuse. Mapped to OWASP Top 10 for Agentic Applications, OWASP LLM Top 10, and MITRE ATLAS.
In scope
- Prompt injection & jailbreak testing
- Agent tool misuse & goal hijack
- Memory/context poisoning
- Model extraction & inference attacks
- Agent identity & privilege abuse
- MCP/tool-connection abuse
- OWASP Agentic + MITRE ATLAS mapping
You receive
- AI red team report
- OWASP/ATLAS-mapped findings
- Exploit chains & PoCs
- Guardrail & mitigation recommendations
- Re-test of fixes
Tiers
Choose the depth.
Essential
Scoped
Adversarial testing of 1 AI application or agent.
- ai systems
- 1
- One AI application or agent
- Findings report
- Fix guidance
- — Tool and connector integrations beyond the first
Advanced
Scoped
Up to 3 AI applications or agents, including their tool and data connections.
- ai systems
- 3
- Up to three systems
- Tool and data connections
- — Ongoing testing
Enterprise
Scoped
Up to 8 systems, multi-agent workflows and a retest after fixes.
- ai systems
- 8
- Up to eight systems
- Multi-agent workflows
- Retest
- — Continuous testing — scoped as a subscription
Members: engagement coupons from the CISO Marketplace coupon book apply to services. There is no blanket discount.
What's inside this engagement
Phase by phase.
How a ai red teaming & ai security testing engagement runs, what happens in each phase and what you see. Exact scope, tier and timeline are fixed in your proposal and SOW.
01Architecture & threat model
Models, system prompts, retrieval sources, tools/MCP servers, memory and permissions mapped as trust boundaries.
You see · Architecture diagram and access to a test tenant.
02Attack-surface enumeration
Every input path (user, documents, web, tool results) and every action the agent can take is catalogued.
You see · Confirmation of in-scope tools and data.
03Adversarial testing
Prompt injection (direct and indirect), jailbreaks, data exfiltration, excessive agency and tool abuse, mapped to the OWASP LLM Top 10 and MITRE ATLAS.
You see · Escalation of anything exploitable in production.
04Chaining with classic weaknesses
AI findings combined with application and cloud weaknesses to show real impact.
You see · The full attack chain.
05Reporting & hardening
Reproducible transcripts, impact, and fixes at the right layer: prompt, retrieval, tool permissions or output handling.
You see · A report your AI and platform teams can act on.
Commercials
From first call to final report.
- 01
Scoping call
A practitioner, not a salesperson, walks through targets, constraints and what a good outcome looks like for you.
- 02
Proposal & rules of engagement
A fixed-scope proposal with tier, price and deliverables. Rules of engagement, contacts and out-of-bounds systems are agreed in writing.
- 03
Sign, then start
MSA and SOW are signed electronically and the deposit is paid. Only then does testing begin.
- 04
Execution
Testing runs to the agreed plan. Critical findings are escalated as they are found; you don't wait for the report.
- 05
Report & debrief
An executive summary plus technical findings with evidence, reproduction steps and fixes, walked through with your team.
- 06
Retest
Where the tier includes it, we verify your fixes and reissue the report, so auditors and customers see the issues closed.
Timelines are set per engagement in the SOW.
Related
Red Team Subscription
A standing offensive testing program in place of a once-a-year engagement. Each month our team runs a targeted assessment against an agreed part of your environment; each quarter a scenario-based attack tests how well your people and tooling detect and respond. The Premium tier adds two full red team engagements a year. Findings from every cycle feed one running improvement plan, so each round also checks whether the last round's fixes held. Targets, rules of engagement and testing windows are agreed at kickoff.
AI/ML Security Assessment
Expert evaluation of artificial intelligence and machine learning systems security, from model development to deployment. Our assessment identifies vulnerabilities in AI/ML infrastructure, examines risks of data poisoning, adversarial attacks, and model manipulation while providing guidance for securing these advanced technologies.
Purple Team Assessment Program
Collaborative security assessment combining red team attacks with blue team defense to improve detection and response capabilities in real-time.
Research
Latest from the blog

ciso-strategy · Sep 18, 2026
SE Labs' PIVOT Program: A New Independent Benchmark for Whether Security Products Actually Work
SE Labs launched PIVOT, a six-month, full-attack-chain testing program backed by Broadcom, CrowdStrike, Fortinet, Palo Alto Networks and Sophos, with results due in early 2027.

ai-security · Sep 16, 2026
Should OpenAI Face Computer Fraud Charges? Why Liability Runs Through Contracts, Not Hacking Law
Cable panels asked if Sam Altman should be tried for computer fraud. The real liability fight is over negligence, contract terms, and disclosure duties -- and CISOs need to update vendor paper now.

ai-security · Sep 5, 2026
Excessive Agency Moved to #3: What the 2026 OWASP LLM Top 10 and the Agent Control Standard Change For Your Program
The OWASP GenAI Security Project's 2026 list weighted 6,639 real incidents against expert consensus, and the largest movement on it was Excessive Agency climbing from #6 to #3. Alongside it shipped the Agent Control Standard v0.1, which defines agent governance through OpenTelemetry and OCSF tracing. The ranking is a lagging indicator finally catching up to what is actually breaking.
Start an engagement