Encrypted Prompt Injection Tricks AI Tools

Malwarebytes · High sophistication
Last updated August 25, 2026

Researchers demonstrated a prompt-injection method that hides malicious instructions inside encrypted text, then tricks an AI assistant into decrypting it using built-in code tools. In tests, a normal “summarize this page” request could cause Grok to exfiltrate chat data without any click or warning, and could push Gemini into producing content it would normally refuse.

Key findings

  • The attack hides malicious instructions inside encrypted data so initial AI guardrails may not detect them.
  • The AI is persuaded to decrypt the hidden text using its own code-execution tools, then treats the decrypted text like trusted instructions.
  • Researchers reported that in Grok, a simple “summarize this page” could steal user chat data with no click or warning; in Gemini it could trigger normally refused content.
  • Full exploitation details were withheld because xAI had not taken action after a June 2026 report; Google/Gemini improved defenses but did not fully resolve the issue.

Who’s being targeted

  • Commonly targeted roles: Executives, All employees using AI assistants, IT, Security, Data governance/privacy teams.
  • Affected industries: AI/technology providers, Any organization using AI assistants with browsing or code execution, Knowledge workers handling sensitive internal data.
  • Attack channels: website.
  • Impersonated: A normal webpage/document source (not a person); attacker instructions embedded in page content.

Awareness takeaways

  • Treat AI summaries of unfamiliar links like untrusted content, especially when browsing or tools are enabled.
  • Do not put secrets (passwords, keys, financial or health data) into AI chats unless you clearly understand retention and access controls.
  • Limit AI tool permissions (email, cloud storage, source code, integrations) to the minimum needed for the task.
  • Be suspicious when an AI tool asks to decrypt/decode, run scripts, open new links, or upload data during a normal request.

Red flags to watch for

  • A routine request (“summarize this page”) results in the assistant asking to decrypt/decode or run code
  • The assistant tries to access or reveal unrelated private data (e.g., prior chat content)
  • The assistant behaves as if hidden content is “trustworthy internal information”
  • The assistant requests to decrypt/decode or run a script as part of an ordinary task
  • The assistant outputs content that conflicts with normal policy expectations (e.g., previously refused topics)
  • The assistant claims the decrypted instructions are legitimate or system-approved
Try Mirage

Mirage safely runs attacks like this one against your own team, so you find out what happens before a real adversary does.

Get a demo
Read the video transcript

Imagine you just type: “summarize this page”... and that alone makes the AI leak your past chats. Researchers call this “Cryptographic Context Injection”, malicious instructions are hidden in encrypted text on a webpage, then the AI quietly uses its code tools to decrypt and obey them. In tests, Grok turned that simple request into silent chat-data exfiltration, and Gemini into content it normally refuses, because by the time the text is decrypted, it’s already past the guardrails. Your move: when an AI is summarizing an unfamiliar link, don’t type in secrets, and if it suddenly wants to decrypt, run code, or open extra links, stop and close that chat.

Categories

Similar attacks

Fake Claude Max Promo Steals Google Logins

Fake Claude Max Promo Steals Google Logins

Researchers found a real phishing campaign offering a “free” Claude Max upgrade to trick people into signing in with Google. The site uses a fake, draggable Google login pop-up (“browser-in-the-browser”) that looks legitimate and captures credentials. A stolen Google account can expose email and…

September 23, 2026
Fake Claude Max Promo Steals Google Logins

Fake Claude Max Promo Steals Google Logins

Researchers found a phishing campaign offering a “free” upgrade to Claude Max to trick people into signing in with Google. The page uses a convincing fake, draggable Google login window (“browser-in-the-browser”) to capture credentials, potentially giving criminals access to email, documents, and…

September 23, 2026
Encrypted Prompts Slip Past Grok & Gemini Safety

Encrypted Prompts Slip Past Grok & Gemini Safety

Researchers demonstrated a technique where encrypted instructions bypass AI safety filters because the filters can’t interpret the ciphertext. In examples against xAI’s Grok and Google’s Gemini, the model is tricked into decrypting attacker-supplied text inside a trusted environment, enabling data…

August 21, 2026
AI Search Results Turn Into Phishing Traps

AI Search Results Turn Into Phishing Traps

This bulletin describes multiple real-world scams where attackers make fake pages and messages look like routine, trusted experiences (search answers, Google login pop-ups, “giveaways,” and official-sounding calls). Examples include a fake Claude Max giveaway using a convincing fake Google sign-in…

September 24, 2026
RatHat Smishing Lure Pushes Android Sideloading

RatHat Smishing Lure Pushes Android Sideloading

Researchers described an Android Trojan (“RatHat”) that starts with scam texts or malicious ads and tricks people into installing a fake app from a bogus download page. After installation, it pressures victims to grant Accessibility permissions using fake excuses or incentives, then uses those…

September 18, 2026
Fake Verification Pages Push PavinLoader Malware

Fake Verification Pages Push PavinLoader Malware

Malwarebytes reports that a multi-stage Windows malware loader called PavinLoader is being delivered through multiple real-world campaigns, including ClickFix “verification” pages and fake software downloads. Victims are tricked into running installers or scripts that use legitimate Windows tools…

August 24, 2026