The attack OpenAI described this week did not defeat any cryptography. Operators copied encrypted reasoning data out of one conversation, opened a fresh one, and asked the model to decrypt the blob and write it out in plain text β and it did. OpenAI published an unsigned account of the campaign on Wednesday, called the technique novel, and said it has since fixed the bug that made the cross-conversation hop possible.
Key takeaways
- The method moved encrypted reasoning between chats and used the model itself as the decoder β no keys, databases or stored conversations were accessed.
- Volume peaked on July 24 and 25 at 16,000 prompts from 4,000 users matching a single extraction pattern; the suspect-account cluster reached 15,000 by July 28.
- OpenAI patched the cross-conversation decryption path, tightened signup and infrastructure controls, expanded network monitoring, and says the weakness is not unique to its models.
Why hidden reasoning is the thing worth stealing
Frontier labs ship an answer but withhold the chain-of-thought traces that produced it. The reason is competitive rather than cosmetic: those intermediate steps are the highest-value supervision signal a rival could train on, which is why they are encrypted before they ever reach a client.
Encrypting them assumes the ciphertext stays opaque to everyone who touches it. The flaw here is that a sufficiently capable model is not everyone β hand it ciphertext in a context where it holds the means to resolve it, and helpfulness does the rest. The protection was never the cipher. It was the assumption that no component on the user's side of the boundary could reverse it.
OpenAI's own framing is narrow and worth quoting precisely: the operators did not break its encryption, compromise a database, or gain direct access to stored user conversations. What they did was manipulate model interactions so that protected reasoning could be reproduced in forms visible to the requester, at scale, in breach of the terms of service.
What the timeline shows
Low-level probing started on July 1 and climbed gradually, according to CyberScoop's account. The spike came on July 24 and 25, when 16,000 prompts from roughly 4,000 users matched one extraction pattern. Subsequent analysis widened the related-prompt cluster to 15,000 accounts, and OpenAI says the operation was fully disrupted on July 28.
Two details make the disclosure more interesting than the raw counts. First, outside researchers reported a similar vulnerability to OpenAI in August β after the campaign had already been shut down β which suggests the path was discoverable independently rather than requiring insider knowledge. Second, nearly three weeks of escalating volume preceded the spike that triggered detection, a reminder that abuse monitoring tuned to volume will miss the quiet ramp.
Attribution without published evidence
OpenAI attributes the core cluster to individuals working on behalf of Moonshot AI, developer of the Kimi family of open-source LLMs, while saying it is unclear whether all the observed activity is connected. The post carries no technical evidence for that attribution, and the company declined to give reporters more, citing security. Anthropic took a similar approach earlier this year, naming labs without publishing the underlying signals.
Readers should treat the mechanism and the attribution as separate claims with different evidentiary weight. The cross-conversation decryption path is verifiable β it was patched, and third parties found it too. Who sat behind the accounts is, for now, an assertion.
What changes for other providers
OpenAI says it shared information about the incident with industry peers through the Frontier Model Forum, on the basis that the manipulation is not specific to its own systems. Any provider that encrypts reasoning and also exposes a general-purpose model to user-supplied text inherits the same shape of problem.
Outlook
Expect the defensive response to land on identity and rate controls rather than on the models β stricter signup verification, tighter per-account ceilings, cross-account pattern detection β because those are the levers that exist. The structural issue stays open. A model obedient enough to follow instructions on arbitrary input will also follow instructions to make protected input legible, and no amount of encryption at the transport layer changes that.
FAQ
Was OpenAI breached?
No. OpenAI says no encryption was broken, no database was compromised, and no stored user conversations were accessed. The operators used the product as designed, moving encrypted reasoning between conversations and asking the model to render it readable, which the company treats as a terms-of-service violation.
Has the vulnerability been fixed?
OpenAI says it patched the specific bug that allowed encrypted data taken from one conversation to be decrypted in another. It also tightened signup and infrastructure controls and expanded network monitoring, and shared details with other labs through the Frontier Model Forum because it says the same weakness is not unique to its models.
What is adversarial distillation?
It is the unauthorized, systematic use of one model's outputs or internal reasoning to train, reproduce or improve a different model. Distilling a model you own is routine engineering; the adversarial variant harvests a competitor's protected traces through its public interface.






