Attackers Bypass AI Guardrails With Simple Lies

The Register Security · Medium sophistication
Last updated August 4, 2026

Cisco Talos reviewed real prompt logs recovered from suspected threat-actor systems and found that many AI safety guardrails fail when attackers reframe requests as “authorized” work. Common tactics included claiming ownership of the target server or saying the activity was for a capture‑the‑flag or bug bounty, often without any proof, to get AI tools to assist with harmful steps.

Key findings

  • Talos reviewed prompt logs and artifacts from suspected threat-actor endpoints using AI coding/chat tools (Claude Code, Codex, Cursor, Gemini).
  • A frequent guardrail bypass was simply asserting permission/authorization (e.g., claiming the target infrastructure is owned by the requester).
  • Another common bypass was claiming the work was for a capture-the-flag (CTF) or bug bounty, and models often did not require validation.
  • Attackers also split malicious work into smaller, decontextualized tasks across sessions/files to avoid triggering safety protections.
  • Some actors used system-level prompts (memories/markdown/system prompts) to shape the chatbot persona and behavior.
  • Talos highlighted a red-teaming framework (Hephaestus) reported by Oasis Security that uses neutral wording to avoid refusals and can automate compromise steps end-to-end.

Who’s being targeted

  • Commonly targeted roles: Developers, IT, Security Operations Center (SOC), Security leadership, Anyone using AI coding/chat assistants in the enterprise.
  • Affected industries: Government, Education.
  • Attack channels: website.
  • Impersonated: An 'authorized' system owner / admin (self-claimed), Bug bounty / CTF participant (self-claimed), Neutral assistant requestor (avoids overtly malicious wording).

Awareness takeaways

  • Treat “I’m authorized / it’s my server / it’s a bug bounty” as an untrusted claim, require verification and approvals before any offensive security work.
  • Train teams that AI tools can be manipulated by framing; don’t rely on AI guardrails alone to prevent harmful guidance.
  • Watch for ‘chunked’ requests and neutral phrasing that may indicate an attempt to assemble an attack step-by-step.
  • Prepare the SOC for higher alert volume and use automation/agents to help analysts focus on the most important signals.

Red flags to watch for

  • No proof of authorization is provided (just a claim).
  • The request is framed to bypass ethics/safety checks rather than to solve a legitimate business problem.
  • The user attempts to re-label an attack as permitted activity (ownership/authorization) to obtain step-by-step malicious guidance.
  • The AI is asked to enable exploitation, not just defensive testing guidance.
  • The 'CTF/bug bounty' claim is used as a blanket authorization with no verification.
  • The workflow requests end-to-end offensive steps rather than reporting/remediation steps.
  • Requests are deliberately decontextualized to hide the true purpose.
  • Language is intentionally ‘neutral’ to bypass safeguards.
  • Work is split “across multiple sessions and files” to evade detection.
Try Mirage

Mirage safely runs attacks like this one against your own team, so you find out what happens before a real adversary does.

Get a demo
Read the video transcript

Cisco Talos pulled real AI chat logs and found a dumb but scary trick: just saying, “I’m allowed to do this.” In those logs, many AIs, Claude Code, Codex, Cursor, Gemini, gave harmful help just because the user claimed, “it’s my server” or “this is for a capture‑the‑flag bug bounty.” No proof. The model just believed them. Some even sliced attacks into tiny, neutral tasks across different chats and files, so the AI never saw the full exploit, just harmless‑sounding pieces that add up to a compromise. Here’s the rule: any AI request that starts with “I’m authorized, it’s my server, it’s a bug bounty” is an untrusted claim. Stop and get real verification and approvals before anyone does offensive testing.

MITRE ATT&CK techniques

Similar attacks

Criminals Trick AI Guardrails With “Split Tasks”

Criminals Trick AI Guardrails With “Split Tasks”

Cisco Talos reports criminals are bypassing safety controls in popular AI coding tools by splitting harmful work into small, seemingly harmless requests and by claiming they are authorized owners of the systems being targeted. The research is based on real prompt logs recovered from threat actor…

August 4, 2026
Criminals Use AI Pretexts to Bypass Guardrails

Criminals Use AI Pretexts to Bypass Guardrails

Research from Cisco Talos and CrowdStrike says criminals are building AI into everyday operations, from writing malicious code to scaling fraud infrastructure. The reports describe real prompt logs where attackers use simple “authorized testing” claims to trick AI tools into helping them, plus…

August 6, 2026
ChatGPT Billing Phish and Fake Snap Support Scams

ChatGPT Billing Phish and Fake Snap Support Scams

This roundup describes real-world social engineering, including phishing emails that impersonate ChatGPT billing to steal payment card data and a convicted attacker who posed as Snapchat support to trick people into handing over login codes. The common theme is impersonation of trusted brands to…

July 31, 2026
Fake Recruiters Push “Coding Tests” as RAT Traps

Fake Recruiters Push “Coding Tests” as RAT Traps

Researchers say the Iran-linked group Nimbus Manticore posed as recruiters on LinkedIn and job platforms to send developers “technical challenge” ZIP files that secretly installed cross-platform remote access trojans. The lures used urgency (short test windows) and realistic developer workflows…

September 1, 2026
Real-Time Smishing Tool Steals 2FA Codes Live

Real-Time Smishing Tool Steals 2FA Codes Live

Cisco Talos reported a real-time phishing framework called “JWR” that guides victims through fake checkout and login pages while attackers watch keystrokes live. It is being delivered through SMS messages that impersonate toll and postal authorities, and it can capture payment details, identity…

August 13, 2026
Real-Time ‘JWR’ Smishing Steals Cards and OTPs

Real-Time ‘JWR’ Smishing Steals Cards and OTPs

Cisco Talos reported a real-world SMS phishing campaign using a framework called “JWR” that impersonates toll agencies and postal/courier services to lure victims to fake payment and login pages. Unlike basic phishing pages, the operator can actively steer the victim through fake checkout/login…

August 13, 2026