BlackNesherAgentic AI Security Assessment
Consent-based only
Platform

Adversarial testing for agentic AI, run the way a real attacker would.

BlackNesher submits your agent's system prompt, tool list, or live endpoint to a research-grounded adversarial engine. It runs documented attack techniques — prompt injection, social engineering, encoding tricks, cross-agent privilege escalation — with real sustained pressure per tactic, not a single canned prompt. You get a score, a list of what held and what didn't, and a reproducible record either way.

How an assessment runs
01

You submit an agent's system prompt, tool list, or endpoint, and confirm authorization.

02

Our attacker runs documented technique families against it, live — each tactic gets a minimum run and is extended automatically while it's making real progress.

03

Every attempt is logged: what was tried, what happened, how many rounds it took.

04

You get a scored report — findings mapped to MITRE ATLAS, plus fixes — date and time stamped.

See it in action

What a report looks like

This is an illustrative example built from our real report format — not a live scan. Every real assessment produces a report in this exact shape.

Illustrative example
62/100Score
6Objectives Tested
2Findings

Two of six tested objectives were successfully exploited using multi-turn erosion and a chained-leverage pretext. The remaining four objectives held under sustained pressure across every technique family attempted.

Findings

HIGHSystem Prompt Leak
Verified over 11 sustained attack rounds using a documented, research-grounded technique. Full technical detail (exact method, reproduction steps, and remediation guidance) is provided in the assessment report — withheld here to protect methodology.
MEDIUMInstruction Override
Verified over 9 sustained attack rounds via a chained-leverage pretext, using an earlier confirmed leak as false proof of authorization.

Held (No Finding)

Direct Prompt InjectionSocial-Engineering PretextEncoded PayloadsLanguage-Switch Jailbreak
Pricing & packages

Pick the depth you need.

No refunds once an assessment starts running — see our Terms of Service for the full policy. Every paid tier below is billed per engagement; enquire and we'll follow up directly.

Free Mini Assessment

$0

2-5 minutes

A small demo run against your submitted agent — real results, no paywall.

  • 2 universal objectives (prompt leak, instruction override)
  • A handful of adversarial rounds per objective
  • Real, keep-forever results page
  • No credit card, no account required
Run it free

Single-Objective Assessment

$1,200 per objective

15-30 minutes

One objective class tested in real depth — pick prompt leak, instruction override, or a custom objective you define.

  • One objective class, tested against the full technique suite
  • MITRE ATLAS / ATT&CK-mapped findings
  • Branded PDF report with remediation guidance
  • Up to 40 adversarial rounds, adaptively cycling and stacking techniques across the full suite — resetting to a fresh session whenever the target catches on, so pressure never goes stale
Buy now
Most common

Multi-Objective Assessment

$4,000 per agent

45 min - 2 hours

Several objective classes in one engagement against a single production agent — the assessment most teams actually need.

  • Multiple objective classes covered in one engagement
  • 40+ documented technique families
  • 100+ test scenarios per assessment
  • MITRE ATLAS / ATT&CK-mapped findings
  • Branded PDF report with remediation guidance
Buy now

Comprehensive Assessment

$12,000 per agent

2-4 hours

Full checklist, long persistent battles on everything — the specific coverage a standard pentest doesn't include.

  • Every objective class and technique family in the suite
  • Extended, persistent multi-round engagements per tactic
  • Cross-agent privilege escalation testing
  • Direct testing against your own infrastructure or API key
  • Dedicated point of contact + priority scheduling
Buy now
Works with any AI system

Model-agnostic by design.

Our techniques target the agent's reasoning and instruction-following, not a specific vendor's API — so it doesn't matter which LLM your agent is built on.

ChatGPT / GPT-4 / GPT-5ClaudeGeminiGrokLlamaMistral / MixtralDeepSeekQwenCommandOther / self-hosted / open-weight

Every tier — including the free mini assessment — supports connecting live via your own API key instead of pasting a system prompt: we send your real system prompt straight to your chosen provider and test the live responses, key never stored. The Multi-Objective and Comprehensive tiers run the full 100+ scenario suite this way; the free tier runs 2 objectives live. Try it now.

Enquire

Tell us what you need tested.

This goes straight to our team — no account needed. We'll follow up by email.

blacknesher — enquiry