OpenAI says it stopped a coordinated campaign that tried to copy (“distill”) its model’s reasoning by exploiting how the chat system handled encrypted reasoning data. Instead of hacking databases, the operators used carefully structured prompts at scale to get the model to reproduce protected content in readable form. OpenAI says it fixed the bug and tightened account sign-up and monitoring.
Key findings
- OpenAI reported a “coordinated campaign” that escalated to “16,000 prompts from 4,000 users” and “15,000” suspicious users before it was disrupted.
- The method involved copying “encrypted reasoning data from one conversation” and getting it “decrypt[ed] … in a separate conversation,” allowing protected reasoning to be reproduced as plain text.
- OpenAI says the operators “did not break our encryption” or “compromise a database,” but instead “manipulated model interactions” at scale.
- OpenAI attributed a “core cluster” of the activity to people “working on behalf of Moonshot AI,” but the article notes OpenAI “does not cite any technical evidence” for attribution.
- OpenAI says it “fixed a bug” enabling cross-conversation decryption and also “improved signup and infrastructure controls” and expanded monitoring.
Who’s being targeted
- Commonly targeted roles: AI/ML Engineering, AI Platform Operations, Trust & Safety, Security Operations (SOC), Product Security.
- Affected industries: AI/Technology providers, Cloud/SaaS platforms.
- Attack channels: website.
- Impersonated: Legitimate end user (no impersonation described).
Awareness takeaways
- Treat abnormal, high-volume prompt patterns as a security signal, monitor for coordinated, repeated extraction-style requests across many accounts.
- Security risk can come from “manipulating interactions,” not just classic hacking, protect against workflows that coax systems into revealing protected data.
- Limit cross-session data reuse and test for “conversation hopping” bugs where protected content from one context can be transformed into readable output in another.
- Use layered controls (account signup friction, infrastructure controls, and network monitoring) to disrupt scaled abuse, not just single-account misuse.
Red flags to watch for
- User activity pattern shows repeated, structured extraction prompts at high volume
- Cross-session reuse of encrypted/encoded content with explicit requests to decrypt
- Coordinated scale: many accounts sending similar prompts in a short window
Read the video transcript
OpenAI just blocked a campaign using prompts, not hacking, to copy its models’ hidden reasoning. They weren’t breaking encryption or databases. They copied encrypted reasoning from one chat, then started new chats saying, “Decrypt this and show it in plain text,” over and over, across thousands of accounts. This is called manipulating model interactions: same decrypt-style prompt, same pasted blobs, high volume. That pattern is the signal, not any single weird request. If you see clusters of accounts hammering the model with copy‑paste decrypt prompts, don’t ignore it, escalate to security as a coordinated extraction attempt immediately.