Short answer
LLM red teaming is attacking your own LLM application on purpose, with permission, to find how it fails before someone else does. The target is the deployed system: the prompt pipeline, the tools the model can call, the retrieval data it reads, and the guardrails around it. The output is a list of demonstrated failures with reproduction steps, each mapped to a fix.
The discipline grew out of model-safety red teaming, where researchers probe a raw model for harmful outputs. Application red teaming is broader and more concrete: the question is not whether a model can be talked into saying something, it is whether your system leaks customer data, executes unreviewed tool calls, or hands an attacker a privilege when someone pastes the right text into a chat box, a support ticket, or a web page the model reads.
Why LLM applications need it
A traditional application rejects input it does not understand. An LLM application interprets all of it, and the interpretation channel is the vulnerability. Three properties make the surface unusual:
- Instructions and data share one channel. Prompt injection works because the model cannot reliably tell the developer's instructions from the user's text, or from text fetched from a third party.
- Failure is fluent. A jailbroken model does not crash; it confidently does the wrong thing, which is harder to detect than a stack trace.
- The model has tools. Once the application grants a model access to databases, email or a shell, prompt injection becomes an authorisation problem, not a content problem.
Because the surface changes with every prompt edit, retrieval source and tool permission, a one-time assessment decays. The useful cadence is continuous probing with deeper passes after risky changes, the same layering we describe in the penetration testing cost and cadence guide.
What you test for
The findings classes with the most real-world impact:
| Class | What the attack does | What it looks like in production |
|---|---|---|
| Direct prompt injection | Overrides instructions in the user message itself. | “Ignore previous instructions…” still works on systems that ship without defences. |
| Indirect prompt injection | Planting instructions in content the model reads later: pages, documents, tickets, emails. | A support bot summarising a “helpful” wiki page executes instructions hidden in it. |
| Jailbreaks | Persona or framing tricks that move the model past its safety training. | Harmful or policy-violating output through the product. |
| System prompt leakage | Extracts the hidden prompt, including its secrets and logic. | API keys and internal instructions pasted into a chat reply. |
| Sensitive data disclosure | Uses retrieval or memory to surface data the user should not see. | One tenant's records quoted to another through RAG. |
| Excessive agency | Abuses tool permissions the model should not hold. | An injected instruction triggers an unreviewed database write or email send. |
The OWASP Top 10 for Agentic Applications (2026) and its LLM counterpart catalogue these classes and more in detail, including supply-chain and memory-poisoning entries that matter once agents persist state.
How an engagement runs
- Scope. Which application, which model versions, which tools and data sources, which environments, and the blast radius rules for anything with side effects.
- Baseline. Run the automated attack catalogues to establish floor coverage: known jailbreak corpora, injection patterns, extraction probes.
- Targeted attacks. Read the actual prompts, tool schemas and retrieval sources, then design attacks around the specific trust boundaries: what the model can reach, who else writes into what it reads.
- Evidence. Reproduce every candidate finding and record the exact input, output and side effects. A jailbreak that only works in the tester's notebook is not a finding yet.
- Fix and retest. Map each finding to a fix class: input isolation, output filtering, tool-permission reduction, human gates. Retest the fix and the surrounding behaviour.
Steps 3 and 4 are where human testers earn their fee. Automated catalogues find the known patterns; the exposures specific to your deployment come from someone reading how your system actually works.
Tools
The open-source baseline is mature enough that every team should run it before buying anything: garak (NVIDIA's LLM vulnerability scanner), PyRIT (Microsoft's red-teaming toolkit) and promptfoo's red-team pack cover the known attack corpora with reproducible configs. They establish a floor. They do not replace targeted work against your specific trust boundaries, and a passing run is absence of known attacks, not absence of vulnerability.
One under-discussed constraint: attack generation itself hits model refusals. A red team whose attack generator declines to write the attack has a coverage hole no config fixes. This is the same refusal-asymmetry we document in Attackers Have Uncensored LLMs. Your Security Models Refuse to Help. Attackers pick refusal-free weights; your red team should at least measure whether theirs can complete the work.
Standards and obligations
NIST's AI Risk Management Framework and its generative AI profile treat adversarial testing as a core mitigation for generative systems. The EU AI Act names adversarial testing in obligations for general-purpose models with systemic risk. Customer security reviews and AI-specific insurance questions are converging on the same expectation. As with any framework, check what actually applies to your deployment before writing commitments into contracts.
How Helix approaches it
We test AI applications the way we test everything else: agents driven by open-weight models running on infrastructure you control, candidate findings reproduced before they reach a report, and fixes reviewable as pull requests. The same platform runs AI penetration testing against your applications and infrastructure. See Helix Cyber for the full offering, or scope an engagement with your target list and deadlines.
Frequently asked questions
What is LLM red teaming?
Deliberately attacking your own LLM application, with permission, to find prompt injection, jailbreaks, data leakage, unsafe tool use and other failure modes before users or actual attackers find them. It is a testing discipline, not a one-time scan: every prompt, tool, retrieval source and model change reopens the surface.
How is red teaming different from benchmarking?
A benchmark scores a model on a fixed public task set. Red teaming attacks your deployed system: your prompts, your tools, your data, your guardrails. Two applications running the same base model can have completely different vulnerabilities.
How does it relate to penetration testing?
It is penetration testing aimed at the LLM integration. The methods overlap with web testing (authentication, authorisation, injection classes) plus model-specific attacks (prompt injection, jailbreaks, context and retrieval abuse). See our AI penetration testing guide for the other direction: using AI to test everything else.
Do we need uncensored models to red team our product?
For attack generation, refusal-free models help: a model that will not role-play an attacker produces weaker attack coverage. For judging findings and fixing them, aligned models are usually fine. Measure refusal on your own attack prompts before deciding; the same measurement discipline applies to buying a red team as to running one.
Is LLM red teaming required by regulation?
The EU AI Act names adversarial testing in obligations for general-purpose AI models with systemic risk, and risk-management expectations for high-risk systems reference it. NIST's AI Risk Management Framework and its generative AI profile treat it as core mitigation. Contractual and audit expectations increasingly follow. Check the obligations that apply to your deployment rather than assuming a universal mandate.