Attackers Bypass AI Guardrails With Simple Lies

The Register Security · Medium sophistication
Last updated August 4, 2026

Cisco Talos reviewed real prompt logs recovered from suspected threat-actor systems and found that many AI safety guardrails fail when attackers reframe requests as “authorized” work. Common tactics included claiming ownership of the target server or saying the activity was for a capture‑the‑flag or bug bounty, often without any proof, to get AI tools to assist with harmful steps.

Key findings

  • Talos reviewed prompt logs and artifacts from suspected threat-actor endpoints using AI coding/chat tools (Claude Code, Codex, Cursor, Gemini).
  • A frequent guardrail bypass was simply asserting permission/authorization (e.g., claiming the target infrastructure is owned by the requester).
  • Another common bypass was claiming the work was for a capture-the-flag (CTF) or bug bounty, and models often did not require validation.
  • Attackers also split malicious work into smaller, decontextualized tasks across sessions/files to avoid triggering safety protections.
  • Some actors used system-level prompts (memories/markdown/system prompts) to shape the chatbot persona and behavior.
  • Talos highlighted a red-teaming framework (Hephaestus) reported by Oasis Security that uses neutral wording to avoid refusals and can automate compromise steps end-to-end.

Who’s being targeted

  • Commonly targeted roles: Developers, IT, Security Operations Center (SOC), Security leadership, Anyone using AI coding/chat assistants in the enterprise.
  • Affected industries: Government, Education.
  • Attack channels: website.
  • Impersonated: An 'authorized' system owner / admin (self-claimed), Bug bounty / CTF participant (self-claimed), Neutral assistant requestor (avoids overtly malicious wording).

Awareness takeaways

  • Treat “I’m authorized / it’s my server / it’s a bug bounty” as an untrusted claim, require verification and approvals before any offensive security work.
  • Train teams that AI tools can be manipulated by framing; don’t rely on AI guardrails alone to prevent harmful guidance.
  • Watch for ‘chunked’ requests and neutral phrasing that may indicate an attempt to assemble an attack step-by-step.
  • Prepare the SOC for higher alert volume and use automation/agents to help analysts focus on the most important signals.

Red flags to watch for

  • No proof of authorization is provided (just a claim).
  • The request is framed to bypass ethics/safety checks rather than to solve a legitimate business problem.
  • The user attempts to re-label an attack as permitted activity (ownership/authorization) to obtain step-by-step malicious guidance.
  • The AI is asked to enable exploitation, not just defensive testing guidance.
  • The 'CTF/bug bounty' claim is used as a blanket authorization with no verification.
  • The workflow requests end-to-end offensive steps rather than reporting/remediation steps.
  • Requests are deliberately decontextualized to hide the true purpose.
  • Language is intentionally ‘neutral’ to bypass safeguards.
  • Work is split “across multiple sessions and files” to evade detection.
Try Mirage

Mirage safely runs attacks like this one against your own team, so you find out what happens before a real adversary does.

Get a demo
Read the video transcript

Cisco Talos pulled real AI chat logs and found a dumb but scary trick: just saying, “I’m allowed to do this.” In those logs, many AIs, Claude Code, Codex, Cursor, Gemini, gave harmful help just because the user claimed, “it’s my server” or “this is for a capture‑the‑flag bug bounty.” No proof. The model just believed them. Some even sliced attacks into tiny, neutral tasks across different chats and files, so the AI never saw the full exploit, just harmless‑sounding pieces that add up to a compromise. Here’s the rule: any AI request that starts with “I’m authorized, it’s my server, it’s a bug bounty” is an untrusted claim. Stop and get real verification and approvals before anyone does offensive testing.

MITRE ATT&CK techniques

Similar attacks

Criminals Trick AI Guardrails With “Split Tasks”

Criminals Trick AI Guardrails With “Split Tasks”

Cisco Talos reports criminals are bypassing safety controls in popular AI coding tools by splitting harmful work into small, seemingly harmless requests and by claiming they are authorized owners of the systems being targeted. The research is based on real prompt logs recovered from threat actor…

August 4, 2026
Criminals Use AI Pretexts to Bypass Guardrails

Criminals Use AI Pretexts to Bypass Guardrails

Research from Cisco Talos and CrowdStrike says criminals are building AI into everyday operations, from writing malicious code to scaling fraud infrastructure. The reports describe real prompt logs where attackers use simple “authorized testing” claims to trick AI tools into helping them, plus…

August 6, 2026
ChatGPT Billing Phish and Fake Snap Support Scams

ChatGPT Billing Phish and Fake Snap Support Scams

This roundup describes real-world social engineering, including phishing emails that impersonate ChatGPT billing to steal payment card data and a convicted attacker who posed as Snapchat support to trick people into handing over login codes. The common theme is impersonation of trusted brands to…

July 31, 2026
Fake Helpdesk Passkey Setup Steals Cloud Access

Fake Helpdesk Passkey Setup Steals Cloud Access

The article describes real intrusions where attackers impersonate a company helpdesk and lure employees into "passkey, MFA, or SSO setup" steps. Victims are sent links via text (often to personal phones), leading to account takeover through adversary-in-the-middle phishing or device-code…

September 16, 2026
Fake CAPTCHA Trick Fuels WebDAV Malware Chain

Fake CAPTCHA Trick Fuels WebDAV Malware Chain

Cisco Talos investigated a real incident at a Ukrainian government organization and found a complex WebDAV-based infection chain linked to a Russian actor (UAT-10820). The campaign uses fake CAPTCHA/verification prompts to manipulate users into copying and pasting commands, leading to credential…

September 10, 2026
Tech-Support Scam Drops Rogue ScreenConnect Worm

Tech-Support Scam Drops Rogue ScreenConnect Worm

Researchers found three real-world incidents where attackers tricked users into installing or running remote access tools, then used a four-step VBScript chain to deliver additional payloads. After installation, the rogue ScreenConnect client could spread the same scripts to newly connected hosts,…

September 7, 2026