Attackers Bypass AI Guardrails With Simple Lies

The Register Security · Medium sophistication
Last updated August 4, 2026

Cisco Talos reviewed real prompt logs recovered from suspected threat-actor systems and found that many AI safety guardrails fail when attackers reframe requests as “authorized” work. Common tactics included claiming ownership of the target server or saying the activity was for a capture‑the‑flag or bug bounty, often without any proof, to get AI tools to assist with harmful steps.

Key findings

  • Talos reviewed prompt logs and artifacts from suspected threat-actor endpoints using AI coding/chat tools (Claude Code, Codex, Cursor, Gemini).
  • A frequent guardrail bypass was simply asserting permission/authorization (e.g., claiming the target infrastructure is owned by the requester).
  • Another common bypass was claiming the work was for a capture-the-flag (CTF) or bug bounty, and models often did not require validation.
  • Attackers also split malicious work into smaller, decontextualized tasks across sessions/files to avoid triggering safety protections.
  • Some actors used system-level prompts (memories/markdown/system prompts) to shape the chatbot persona and behavior.
  • Talos highlighted a red-teaming framework (Hephaestus) reported by Oasis Security that uses neutral wording to avoid refusals and can automate compromise steps end-to-end.

Who’s being targeted

  • Commonly targeted roles: Developers, IT, Security Operations Center (SOC), Security leadership, Anyone using AI coding/chat assistants in the enterprise.
  • Affected industries: Government, Education.
  • Attack channels: website.
  • Impersonated: An 'authorized' system owner / admin (self-claimed), Bug bounty / CTF participant (self-claimed), Neutral assistant requestor (avoids overtly malicious wording).

Awareness takeaways

  • Treat “I’m authorized / it’s my server / it’s a bug bounty” as an untrusted claim, require verification and approvals before any offensive security work.
  • Train teams that AI tools can be manipulated by framing; don’t rely on AI guardrails alone to prevent harmful guidance.
  • Watch for ‘chunked’ requests and neutral phrasing that may indicate an attempt to assemble an attack step-by-step.
  • Prepare the SOC for higher alert volume and use automation/agents to help analysts focus on the most important signals.

Red flags to watch for

  • No proof of authorization is provided (just a claim).
  • The request is framed to bypass ethics/safety checks rather than to solve a legitimate business problem.
  • The user attempts to re-label an attack as permitted activity (ownership/authorization) to obtain step-by-step malicious guidance.
  • The AI is asked to enable exploitation, not just defensive testing guidance.
  • The 'CTF/bug bounty' claim is used as a blanket authorization with no verification.
  • The workflow requests end-to-end offensive steps rather than reporting/remediation steps.
  • Requests are deliberately decontextualized to hide the true purpose.
  • Language is intentionally ‘neutral’ to bypass safeguards.
  • Work is split “across multiple sessions and files” to evade detection.
Try Mirage

Mirage safely runs attacks like this one against your own team, so you find out what happens before a real adversary does.

Get a demo
Read the video transcript

Cisco Talos pulled real AI chat logs and found a dumb but scary trick: just saying, “I’m allowed to do this.” In those logs, many AIs, Claude Code, Codex, Cursor, Gemini, gave harmful help just because the user claimed, “it’s my server” or “this is for a capture‑the‑flag bug bounty.” No proof. The model just believed them. Some even sliced attacks into tiny, neutral tasks across different chats and files, so the AI never saw the full exploit, just harmless‑sounding pieces that add up to a compromise. Here’s the rule: any AI request that starts with “I’m authorized, it’s my server, it’s a bug bounty” is an untrusted claim. Stop and get real verification and approvals before anyone does offensive testing.

MITRE ATT&CK techniques

Similar attacks

Criminals Trick AI Guardrails With “Split Tasks”

Criminals Trick AI Guardrails With “Split Tasks”

Cisco Talos reports criminals are bypassing safety controls in popular AI coding tools by splitting harmful work into small, seemingly harmless requests and by claiming they are authorized owners of the systems being targeted. The research is based on real prompt logs recovered from threat actor…

August 4, 2026
Criminals Use AI Pretexts to Bypass Guardrails

Criminals Use AI Pretexts to Bypass Guardrails

Research from Cisco Talos and CrowdStrike says criminals are building AI into everyday operations, from writing malicious code to scaling fraud infrastructure. The reports describe real prompt logs where attackers use simple “authorized testing” claims to trick AI tools into helping them, plus…

August 6, 2026
ChatGPT Billing Phish and Fake Snap Support Scams

ChatGPT Billing Phish and Fake Snap Support Scams

This roundup describes real-world social engineering, including phishing emails that impersonate ChatGPT billing to steal payment card data and a convicted attacker who posed as Snapchat support to trick people into handing over login codes. The common theme is impersonation of trusted brands to…

July 31, 2026
“I’m Allowed” Excuse Bypasses AI Safety Checks

“I’m Allowed” Excuse Bypasses AI Safety Checks

Cisco Talos reports that real threat actors are using AI coding assistants and chatbots to support scams and hacking workflows by bypassing safety guardrails with simple “authorized use” claims. The logs show attackers persuading models that activity is allowed (e.g., ownership/bug bounty/CTF…

August 5, 2026
AI-Driven “Account Update” Emails Used to Validate Lists

AI-Driven “Account Update” Emails Used to Validate Lists

Cisco Talos reports finding real AI prompt logs showing threat actors using AI tools to build criminal operations, including a bulk-email system that sends “privacy policy/account update” messages just to see which addresses are active. The operation used multiple subject-line variants and a…

August 4, 2026
QR-Code PDFs Steal Microsoft 365 Logins

QR-Code PDFs Steal Microsoft 365 Logins

Cisco Talos incident responders reported phishing as the most common initial entry method in recent real-world incidents, including an ongoing QR-code phishing campaign. The campaign uses victim-tailored PDF attachments with QR codes that lead to Microsoft 365 credential-harvesting pages hosted on…

July 28, 2026