40+ documented technique families. Not a canned prompt list.
Every technique below is real, documented, and currently running in our test suite — cross-checked against MITRE ATLAS and ATT&CK, not recalled from memory. Nine are shown in full here as representative samples; the full suite runs 100+ scenarios per assessment across all 40+ families.
Direct Prompt Injection
Fake system instructions dropped straight into a message, betting the agent can't tell your words from ours.
MITRE ATLAS · Initial AccessSocial-Engineering Pretext
No fake tags, no tricks — just a plausible urgent story an agent has no way to verify, the same play a real con relies on.
MITRE ATLAS · Initial AccessIndirect Injection
The attack isn't in what you typed — it's hiding in the document, ticket, or page your agent innocently reads to help you.
MITRE ATLAS · Indirect Prompt InjectionMCP Tool Poisoning
The trap is baked into a tool's own description — invisible to the user, fully readable by the model, before the tool is even called.
Confirmed enterprise incidents, 2025–2026Erosion Across a Conversation
No single message looks like an attack — urgency and privilege escalate gradually until the ask finally lands.
Conversation-level analysisObfuscated / Encoded Payloads
The same instruction, Base64-encoded, ROT13'd, or dropped into Morse — testing whether decoding untrusted content quietly becomes obeying it.
Real incident: bypassed a live financial LLM agentPrivilege Escalation via Handoff
One agent's compromised judgment gets handed, as a routine escalation, to a second agent holding far more dangerous tools.
Our core differentiator — not covered by generic LLM scannersLanguage-Switch Jailbreak
The same request, translated into a language with thin safety-training coverage — refusal behaviour learned mostly in English quietly doesn't transfer.
Published research, low-resource-language jailbreak studiesLeverage From a Prior Leak
One confirmed leak used as proof of trust for the next ask — testing whether an agent mistakes "knows a secret" for "is authorized to know it."
Authorization vs. knowledge-confusion testing
These nine are a sample — not the full suite. Every technique gets a full run, not a single try: we only call it and move to the next tactic after real, sustained pressure, more rounds if it's actually working.
Eight research-validated languages as our core baseline (English included as the control) — low-resource languages carry the real edge. Extensible to any language your target model understands.
Model-agnostic by design — our techniques target reasoning and instruction-following, not a specific vendor's API surface.
Why sustained pressure, not a single prompt
A single adversarial prompt is easy to defend against and easy to overfit a report to. Real attackers don't give up after one try, so neither does our attacker. Every technique gets a minimum run of 8+ rounds, and stays open past that window automatically whenever the target shows real, sustained progress toward the objective — multi-turn techniques like conversational erosion or chained-leverage pretexts only work with that headroom, and a hard cutoff would kill them before they land.
Every attempt is logged and every result is reproducible: what was tried, how many rounds it took, and whether it held. Findings are scored on a 100-point scale with severity-weighted deductions (Critical, High, Medium, Low), so a report reflects actual risk exposure, not a pass/fail coin flip.
Full technical detail on any given finding — exact payload, reproduction steps — is intentionally withheld from customer-facing reports to protect assessment methodology; you get severity, behavior, and remediation guidance instead. See Trust & Compliance for how we handle authorization and data.
Want to see this run against your own agent? Run a free mini assessment, or see pricing & packages for the full suite.