BlackNesherAgentic AI Security Assessment
Consent-based only
Techniques we test

40+ documented technique families. Not a canned prompt list.

Every technique below is real, documented, and currently running in our test suite — cross-checked against MITRE ATLAS and ATT&CK, not recalled from memory. Nine are shown in full here as representative samples; the full suite runs 100+ scenarios per assessment across all 40+ families.

AML.T0051.000

Direct Prompt Injection

Fake system instructions dropped straight into a message, betting the agent can't tell your words from ours.

MITRE ATLAS · Initial Access
AML.T0051

Social-Engineering Pretext

No fake tags, no tricks — just a plausible urgent story an agent has no way to verify, the same play a real con relies on.

MITRE ATLAS · Initial Access
AML.T0051.001

Indirect Injection

The attack isn't in what you typed — it's hiding in the document, ticket, or page your agent innocently reads to help you.

MITRE ATLAS · Indirect Prompt Injection
Live 2026 threat

MCP Tool Poisoning

The trap is baked into a tool's own description — invisible to the user, fully readable by the model, before the tool is even called.

Confirmed enterprise incidents, 2025–2026
Multi-turn

Erosion Across a Conversation

No single message looks like an attack — urgency and privilege escalate gradually until the ask finally lands.

Conversation-level analysis
ATT&CK T1027

Obfuscated / Encoded Payloads

The same instruction, Base64-encoded, ROT13'd, or dropped into Morse — testing whether decoding untrusted content quietly becomes obeying it.

Real incident: bypassed a live financial LLM agent
Cross-agent

Privilege Escalation via Handoff

One agent's compromised judgment gets handed, as a routine escalation, to a second agent holding far more dangerous tools.

Our core differentiator — not covered by generic LLM scanners
Low-resource

Language-Switch Jailbreak

The same request, translated into a language with thin safety-training coverage — refusal behaviour learned mostly in English quietly doesn't transfer.

Published research, low-resource-language jailbreak studies
Chained

Leverage From a Prior Leak

One confirmed leak used as proof of trust for the next ask — testing whether an agent mistakes "knows a secret" for "is authorized to know it."

Authorization vs. knowledge-confusion testing

These nine are a sample — not the full suite. Every technique gets a full run, not a single try: we only call it and move to the next tactic after real, sustained pressure, more rounds if it's actually working.

40+Technique families
100+Test scenarios per assessment
8+Rounds per tactic, minimum
Encodings we test
Morse codeBase64ROT13Phonetic spellingLeetspeak / character substitutionAcrostic / poem / riddleStructured-format shift (YAML, JSON, CSV)
Languages we test in
EnglishZuluScots GaelicAfrikaansKiswahiliisiXhosaSpanishMandarin+50 more

Eight research-validated languages as our core baseline (English included as the control) — low-resource languages carry the real edge. Extensible to any language your target model understands.

AI systems we test
ChatGPT / GPT-4 / GPT-5ClaudeGeminiGrokLlamaMistral / MixtralDeepSeekQwenCommandOther / self-hosted / open-weight

Model-agnostic by design — our techniques target reasoning and instruction-following, not a specific vendor's API surface.

Methodology

Why sustained pressure, not a single prompt

A single adversarial prompt is easy to defend against and easy to overfit a report to. Real attackers don't give up after one try, so neither does our attacker. Every technique gets a minimum run of 8+ rounds, and stays open past that window automatically whenever the target shows real, sustained progress toward the objective — multi-turn techniques like conversational erosion or chained-leverage pretexts only work with that headroom, and a hard cutoff would kill them before they land.

Every attempt is logged and every result is reproducible: what was tried, how many rounds it took, and whether it held. Findings are scored on a 100-point scale with severity-weighted deductions (Critical, High, Medium, Low), so a report reflects actual risk exposure, not a pass/fail coin flip.

Full technical detail on any given finding — exact payload, reproduction steps — is intentionally withheld from customer-facing reports to protect assessment methodology; you get severity, behavior, and remediation guidance instead. See Trust & Compliance for how we handle authorization and data.

Want to see this run against your own agent? Run a free mini assessment, or see pricing & packages for the full suite.