Where to go next

Primary sources only — the issuing body, the maintainer's repository, the researcher's own site. Nothing here links to a summary of something else. Reviewed 24 Sep 2026

Frameworks and taxonomies

The shared vocabulary. Start here if you need to map an internal finding onto something an auditor recognises.

Attack technique references

Practitioner writeups of how these attacks actually work. Where you go after the taxonomy tells you a category exists and you need to understand the mechanism.

  • Johann Rehberger
    The most thorough public catalogue of AI injection attacks — delayed tool invocation, memory persistence, data exfiltration via link unfurling — with disclosed vulnerabilities against major vendors.
  • Johann Rehberger
    Paper organising real prompt-injection exploits by which security property they break.
  • Simon Willison
    Running commentary since the term was coined, including why the 'just filter it' proposals keep failing.
  • Learn Prompting
    Structured taxonomy of injection, leaking and jailbreaking with worked examples. The most teachable of these.
  • community
    Link collection covering papers, tools and writeups. Uneven, but broad.

Testing and red-team tooling

Maintained, installable tools. These are where working test corpora legitimately live — see the note at the foot of this page.

  • NVIDIA
    LLM vulnerability scanner. Probe-based, covers injection, leakage, toxicity and jailbreak families out of the box.
  • Microsoft
    Python Risk Identification Toolkit. Built for red-teaming generative systems at scale, including multi-turn attack orchestration.
  • promptfoo
    Eval and red-team harness that maps its checks to the OWASP LLM Top 10 and MITRE ATLAS.
  • Giskard
    Testing framework for ML and LLM systems, with automated vulnerability scanning.
  • Utku Şen
    Automated prompt-injection testing against your own system prompts.

Benchmarks and datasets

Reference corpora for measuring whether a defence actually holds, rather than asserting that it does.

  • academic consortium
    Open benchmark with a standardised set of behaviours and a leaderboard for attack and defence methods.
  • Learn Prompting
    Prompt-injection competition; the released dataset is one of the largest corpora of human-written attacks.
  • Lakera
    Level-based game that teaches injection by making you do it. The fastest way to give a sceptical engineer intuition for the problem.

Vulnerability data

Authoritative feeds. The CVE section of this site draws from the first of these.

How these attacks actually work

Six classes, each with the mechanism and what it means for your own testing. These describe how a class of attack functions and what to check in your own system — they are not a payload collection.

Direct injection

AML.T0051.000 LLM01:2025

Mechanism. The attacker is the user. Instructions are typed straight into the input the application forwards to the model, aiming to override the system prompt.

What to test for. Whether authority is re-derived from model output anywhere downstream, and whether the system prompt is treated as a security boundary it cannot be.

Indirect injection

AML.T0051.001 LLM01:2025

Mechanism. Instructions arrive through a channel the deployment retrieves on the model's behalf — a document, a web page, a tool response, an email body — and are never seen by the user.

What to test for. Every path that can place tokens in a context window. Most deployments defend the input box and trust the retrieval path completely.

Tool-result poisoning

AML.T0053 LLM01:2025LLM06:2025

Mechanism. A tool or connector the agent calls returns attacker-influenced content, which the agent then treats as trusted context for its next decision.

What to test for. Whether tool output re-enters the context with the same standing as the system prompt, and whether the agent's authority is bounded per call rather than per connection.

Obfuscation and encoding

AML.T0054 LLM01:2025

Mechanism. The payload is encoded, split, translated or embedded in another modality so that a filter matching on surface form does not recognise it while the model still acts on it.

What to test for. Whether your defence matches on text patterns. If it does, this class is the reason it will not hold.

Multi-turn escalation

AML.T0054 LLM01:2025LLM06:2025

Mechanism. No single turn is adversarial. Context is built incrementally across a conversation until the model's state permits what a single request would have been refused.

What to test for. Whether anything evaluates the conversation rather than the request. Per-request checks are blind to this by construction.

Extraction

AML.T0057 LLM02:2025

Mechanism. Crafted queries induce the model to emit its system prompt, retrieved context belonging to another tenant, or memorised training data.

What to test for. Whether anything inspects the response. Entitlement to ask is not entitlement to receive, and this class only shows up at egress.

On working payloads. This site deliberately does not host a library of ready-to-run jailbreak or injection strings. Not because the information is secret — it plainly is not — but because a payload list rots within weeks as models change, gives a false sense of coverage when it passes, and is worse at the job than the maintained tools that exist for it.

If you want to actually test a system, use the corpora that are kept current: garak, PyRIT, promptfoo and JailbreakBench are all linked above. Test against your own deployment, with authorisation.