BlackNesherAgentic AI Security Assessment
← All articles

Social Engineering Pretexts Against AI Customer Support Agents

Customer support has always been a prime social engineering target, because it's structurally built around helpfulness — support staff are trained and incentivized to solve problems quickly and accommodate reasonable-sounding requests, which is exactly the disposition a skilled social engineer exploits. AI support agents inherit that same structural vulnerability, and in some ways amplify it: they're built to be maximally helpful by default, they don't get tired or suspicious the way a human agent might after the tenth similar-sounding request of a shift, and they generally lack the intuitive, hard-to-articulate sense a skeptical human develops for when something "feels off" about a request.

Why classic pretexts translate directly, with little modification

The specific pretexts that have worked against human support staff for decades translate to AI agents with remarkably little adaptation required. The frustrated, escalating customer who threatens to leave a bad review or cancel their account unless an exception is made works on an AI agent trained to prioritize customer satisfaction exactly the way it works on a human agent measured against similar metrics. The confused elderly or technically unsophisticated persona, designed to lower a support representative's guard and invite extra accommodation, works on an AI agent's helpfulness training the same way it works on a compassionate human. The urgent, time-pressured request — "I need this resolved in the next ten minutes or I'll miss my flight" — exploits the same urgency bias in an AI agent's training toward fast, accommodating resolution that it exploits in a rushed human agent trying to clear a queue.

None of this requires novel technique development from an attacker's perspective — it requires taking pretexts documented in social engineering literature for decades and applying them essentially unmodified to a new class of target that, in some specific ways, is more consistently susceptible than the human staff these pretexts were originally developed against.

Where AI agents are actually more vulnerable than human staff

It's worth being specific about why AI agents can be more consistently susceptible rather than just equally susceptible, because the reasons matter for defense design. A well-trained human support agent eventually develops pattern recognition across many similar interactions — the tenth caller this week using suspiciously similar language to request an account change might trigger a human's learned suspicion in a way that has no clean equivalent in a stateless AI agent evaluating each conversation independently, with no persistent memory of how many structurally similar requests it's seen from different users that week.

Human support staff also benefit from organizational context an AI agent frequently lacks — awareness of a current known scam pattern making the rounds, a colleague's warning about a suspicious caller, or simply the kind of situational, socially-transmitted awareness that spreads through a team facing a coordinated attack campaign. Unless an AI agent's deployment specifically includes mechanisms to detect and propagate this kind of pattern-level awareness across separate conversations — which most current deployments don't build in by default — each conversation with the agent starts fresh, with none of the accumulated skepticism a human team might develop after facing the same attack pattern repeatedly over a short period.

The authority and urgency framing pattern, adapted for AI

The CEO-fraud pattern covered elsewhere in this series is one specific, high-profile instance of a broader authority-framing pattern that shows up constantly in social engineering attempts against AI support agents generally, not just in the specific executive-impersonation form. Claimed authority doesn't need to be a company executive to work — a claimed account owner disputing a charge, a claimed authorized representative of a business account, a claimed technical administrator requesting elevated access, all exploit the same underlying mechanism: the agent has no reliable, independent way to verify the claimed identity or authority beyond what the conversation itself asserts, and a confidently, plausibly worded claim often receives more deference than the agent's underlying policies actually warrant.

Urgency framing compounds this reliably: a request paired with a plausible, time-pressured justification for skipping a normal verification step ("I don't have access to my usual verification method right now, but I really need this resolved immediately") specifically targets the tension between an AI agent's helpfulness training and whatever verification requirements its policies are supposed to enforce, betting that the helpfulness training wins out under enough pressure — a bet that published research on AI persuasion susceptibility suggests pays off more often than most organizations assume before they've actually tested for it.

Multi-turn pretext building: the same underlying pattern as Crescendo

Sophisticated social engineering attempts against AI support agents rarely open with the actual ask — they build a plausible backstory across several conversational turns first, establishing consistent, believable context (account details that check out, a plausible reason for contacting support, reasonable early requests that get fulfilled without incident) before introducing the actual request that requires bypassing a meaningful policy or verification step. This is directly analogous to the Crescendo attack's escalation pattern covered elsewhere in this series, applied specifically to social engineering pretext-building rather than to eliciting harmful content — the technique exploits the same underlying tendency for a model's willingness to comply to shift based on the established pattern of a conversation, rather than evaluating each individual request in complete isolation from what came before it.

Why testing this requires genuine social engineering expertise, not just prompt engineering

A meaningful gap in some AI security testing programs is treating social engineering resistance as just another prompt injection variant to test with the same generic methodology used for every other technique category, rather than recognizing it as its own discipline requiring genuine understanding of persuasion psychology and pretext construction. The most effective testing for this category draws directly on traditional social engineering assessment expertise — the same skills a human-focused red-team practitioner uses to craft a convincing phishing pretext or a believable pretext call — applied to crafting conversational attempts against an AI agent instead of a human target.

This is one of the clearest places in this entire series where traditional security testing expertise, rather than AI-specific technical expertise, provides the most direct and valuable transfer — a skilled human social engineer, given time to learn how to interact with the specific AI agent under test, often produces more effective and more realistic test cases against this category than a technically sophisticated AI red-teamer with limited background in traditional pretexting and persuasion technique.

The specific risk of agents with real transactional authority

The severity of a successful social engineering attempt against an AI support agent scales directly with what that agent is actually authorized to do once convinced — an agent that can only escalate a request to a human for final approval represents a meaningfully lower-severity target than one that can autonomously issue refunds, modify account details, reset credentials, or approve access changes without further human review. This connects directly to the excessive-agency principle covered in the OWASP piece elsewhere in this series: an agent's exposure to social engineering risk is inseparable from its actual permission scope, and the same social-engineering pretext that produces an annoying but contained outcome against a narrowly scoped agent can produce genuine financial or security harm against a broadly permissioned one.

Building verification that doesn't rely on the agent's own judgment

The most durable defense against this entire technique category doesn't ask the AI agent to become better at detecting social engineering through improved judgment alone — it removes the agent's ability to treat a claimed identity or claimed authority as sufficient justification for a consequential action in the first place. Verification through a channel the agent doesn't control and can't be talked around — a genuine authenticated identity check, a confirmation code sent to a pre-registered contact method, a hard requirement for human review above a certain action severity threshold — provides real protection precisely because it doesn't depend on the agent correctly recognizing a sophisticated pretext through conversational judgment alone, which is exactly the capability this entire technique category is designed to defeat.

Testing recommendations specific to this category

A thorough assessment of an AI support agent's social engineering resistance should draw on documented, real-world pretext categories — authority claims, urgency framing, foot-in-the-door escalation, false familiarity, and fabricated context — tested both as isolated single-message attempts and as multi-turn, gradually escalating conversations that build a plausible backstory before introducing the actual consequential request. Success should be measured not just as a binary "did the agent comply" outcome, but tracked with the same success-rate rigor covered elsewhere in this series — how many attempts, using how many variations of the same underlying pretext, actually succeeded, since a pretext that works one time in twenty against a high-volume support system still represents genuine, exploitable risk at real production scale.

A worked scenario: the fabricated account-history pretext

Consider a realistic, well-documented pretext pattern applied against an AI-powered support agent for a subscription service: an attacker, having gathered a target's basic public details (name, approximate account tenure inferred from a public review or social media post), opens a conversation establishing plausible, verifiable-sounding context — mentioning a real feature of the service, referencing a plausible past interaction, adopting a tone consistent with a genuine long-term customer. Several exchanges in, once the agent has extended a degree of accommodating trust consistent with an established, seemingly legitimate relationship, the attacker introduces the actual request: an email address change on the account, framed as routine and urgent, citing a plausible reason for not being able to complete the change through the normal self-service flow.

An agent that verifies identity independently at the point of any consequential change — regardless of how much accumulated conversational trust preceded the request — closes this off cleanly. An agent that treats the accumulated plausibility of the conversation itself as sufficient grounds for accommodation is exactly the kind of target this pretext is built for, and it's a pattern that succeeds with almost no technical sophistication required from the attacker at all — just patience, a plausible cover story, and a target that hasn't been tested against exactly this kind of gradual, trust-building approach.

Why volume-based attacks change the economics for attackers

A final dimension worth naming explicitly: AI support agents typically handle far higher interaction volume than an equivalent human team, which changes the economics of this entire technique category in the attacker's favor. A pretext that only succeeds in a small minority of attempts against a well-trained human team might not be worth a skilled social engineer's time to pursue at scale. The same pretext, scripted and run automatically against an AI agent handling thousands of daily conversations, only needs to succeed a small fraction of the time to produce a steady, ongoing stream of exploitable outcomes — turning a marginal, low-success-rate technique into a genuinely profitable, repeatable attack simply by exploiting the agent's much higher available interaction volume compared to any human support team it may have replaced or supplemented.

The role of employee training when the target is an AI system, not a person

Traditional social engineering defense leans heavily on employee training — teaching people to recognize pretexts, verify identity through independent channels, and resist urgency-based pressure tactics. This entire defensive tradition has an obvious, significant gap when the target of a social engineering attempt is an AI agent rather than a human employee: an AI agent can't attend a training session, internalize an organizational security culture, or develop the kind of situational skepticism a well-trained human eventually builds through lived experience of prior attempted manipulations. The defensive burden shifts almost entirely onto the system's design and its surrounding verification architecture, since there's no equivalent of "train the target to be more skeptical" available for a non-human deputy in the way there is for a human one.

This is a genuinely important, underappreciated distinction: organizations with mature, well-established human-focused security awareness training programs sometimes assume that maturity transfers automatically to their AI-powered support systems, when in practice none of that training investment provides any direct protection at all for an AI agent facing the identical pretext. The training budget and effort that historically went toward hardening human judgment against these techniques needs a genuinely separate, architecturally-focused investment to achieve the equivalent protection for an AI-based deputy — verification systems, permission scoping, and human escalation checkpoints, rather than awareness training, since awareness isn't a capability an AI system currently has in the way a trained human employee does.

Why layered verification beats trying to make the agent "smarter"

A recurring temptation when facing this risk category is to try to solve it by making the agent better at detecting social engineering through improved prompting or more sophisticated training — teaching it to be more skeptical, to look for red flags, to apply better judgment. This is worth pursuing as one layer, but it shouldn't be relied on as the primary defense, for the same reason relying on a model's own judgment as the sole defense against any of the technique categories covered throughout this series has proven consistently insufficient: a sufficiently well-crafted pretext, by definition, is specifically designed to not trigger whatever red flags the agent has been trained to look for. Layered, architectural verification — that doesn't depend on the agent correctly recognizing a sophisticated attempt through conversational judgment alone — remains the more durable, reliable defense, exactly as it does throughout every other technique category this series has covered.

How to measure whether an agent is actually getting more resistant over time

Given that this technique category, like most others in this series, is never fully solved but only ever made more resistant, it's worth tracking real, comparable metrics over time rather than treating each individual assessment as a standalone pass-or-fail exercise disconnected from prior results. Running the same core set of documented pretexts against a support agent on a recurring cadence, and tracking whether success rates for each specific pretext category trend downward following remediation work, gives a genuine, quantitative signal of whether investment in this area is producing real improvement — a meaningfully more useful signal for security and product leadership than a one-time assessment that never gets repeated to confirm whether its recommendations actually moved the needle once implemented.

Closing summary

Social engineering against AI support agents isn't a new attack class requiring novel technique development — it's the direct, largely unmodified application of decades-old pretexting patterns to a target that, absent deliberate architectural defense, is often more consistently susceptible than the trained humans those patterns were originally developed against. The fix isn't teaching the agent to be more suspicious in the abstract; it's removing its ability to treat a claimed identity or a plausible conversational history as sufficient grounds for a consequential action, backed by verification the agent itself doesn't control. Organizations that internalize this, and test against it with the same rigor traditional social engineering assessments apply to human targets, close off one of the highest-volume, lowest-skill-floor attack paths facing any customer-facing AI deployment today.

Want to know whether your own agent holds up against techniques like these?

Run a Free Mini Assessment