AI Newsway

A Cross-Chat Bug Let OpenAI's Model Decode Its Own Reasoning

OpenAI says operators pasted encrypted reasoning blobs from one conversation into another and asked the model to transcribe them. The flaw is patched; the company says the same path exists elsewhere.

|5 min read0
AI Summary
OpenAI disclosed on September 30, 2026 that operators extracted its models' protected reasoning by copying encrypted traces from one conversation and asking the model to transcribe them in another. Activity began July 1 and peaked at 16,000 prompts from 4,000 users on July 24-25, spanning 15,000 accounts by July 28. OpenAI patched the cross-conversation path, says no encryption or database was breached, and shared details through the Frontier Model Forum.
The office building at 1515 Third Street in San Francisco's Mission Bay neighborhood, which has housed OpenAI's headquarters.
The office building at 1515 Third Street in San Francisco's Mission Bay neighborhood, which has housed OpenAI's headquarters.

The attack OpenAI described this week did not defeat any cryptography. Operators copied encrypted reasoning data out of one conversation, opened a fresh one, and asked the model to decrypt the blob and write it out in plain text β€” and it did. OpenAI published an unsigned account of the campaign on Wednesday, called the technique novel, and said it has since fixed the bug that made the cross-conversation hop possible.

Key takeaways

  • The method moved encrypted reasoning between chats and used the model itself as the decoder β€” no keys, databases or stored conversations were accessed.
  • Volume peaked on July 24 and 25 at 16,000 prompts from 4,000 users matching a single extraction pattern; the suspect-account cluster reached 15,000 by July 28.
  • OpenAI patched the cross-conversation decryption path, tightened signup and infrastructure controls, expanded network monitoring, and says the weakness is not unique to its models.

Why hidden reasoning is the thing worth stealing

Frontier labs ship an answer but withhold the chain-of-thought traces that produced it. The reason is competitive rather than cosmetic: those intermediate steps are the highest-value supervision signal a rival could train on, which is why they are encrypted before they ever reach a client.

Encrypting them assumes the ciphertext stays opaque to everyone who touches it. The flaw here is that a sufficiently capable model is not everyone β€” hand it ciphertext in a context where it holds the means to resolve it, and helpfulness does the rest. The protection was never the cipher. It was the assumption that no component on the user's side of the boundary could reverse it.

OpenAI's own framing is narrow and worth quoting precisely: the operators did not break its encryption, compromise a database, or gain direct access to stored user conversations. What they did was manipulate model interactions so that protected reasoning could be reproduced in forms visible to the requester, at scale, in breach of the terms of service.

What the timeline shows

Low-level probing started on July 1 and climbed gradually, according to CyberScoop's account. The spike came on July 24 and 25, when 16,000 prompts from roughly 4,000 users matched one extraction pattern. Subsequent analysis widened the related-prompt cluster to 15,000 accounts, and OpenAI says the operation was fully disrupted on July 28.

Two details make the disclosure more interesting than the raw counts. First, outside researchers reported a similar vulnerability to OpenAI in August β€” after the campaign had already been shut down β€” which suggests the path was discoverable independently rather than requiring insider knowledge. Second, nearly three weeks of escalating volume preceded the spike that triggered detection, a reminder that abuse monitoring tuned to volume will miss the quiet ramp.

Attribution without published evidence

OpenAI attributes the core cluster to individuals working on behalf of Moonshot AI, developer of the Kimi family of open-source LLMs, while saying it is unclear whether all the observed activity is connected. The post carries no technical evidence for that attribution, and the company declined to give reporters more, citing security. Anthropic took a similar approach earlier this year, naming labs without publishing the underlying signals.

Readers should treat the mechanism and the attribution as separate claims with different evidentiary weight. The cross-conversation decryption path is verifiable β€” it was patched, and third parties found it too. Who sat behind the accounts is, for now, an assertion.

What changes for other providers

OpenAI says it shared information about the incident with industry peers through the Frontier Model Forum, on the basis that the manipulation is not specific to its own systems. Any provider that encrypts reasoning and also exposes a general-purpose model to user-supplied text inherits the same shape of problem.

Outlook

Expect the defensive response to land on identity and rate controls rather than on the models β€” stricter signup verification, tighter per-account ceilings, cross-account pattern detection β€” because those are the levers that exist. The structural issue stays open. A model obedient enough to follow instructions on arbitrary input will also follow instructions to make protected input legible, and no amount of encryption at the transport layer changes that.

FAQ

Was OpenAI breached?

No. OpenAI says no encryption was broken, no database was compromised, and no stored user conversations were accessed. The operators used the product as designed, moving encrypted reasoning between conversations and asking the model to render it readable, which the company treats as a terms-of-service violation.

Has the vulnerability been fixed?

OpenAI says it patched the specific bug that allowed encrypted data taken from one conversation to be decrypted in another. It also tightened signup and infrastructure controls and expanded network monitoring, and shared details with other labs through the Frontier Model Forum because it says the same weakness is not unique to its models.

What is adversarial distillation?

It is the unauthorized, systematic use of one model's outputs or internal reasoning to train, reproduce or improve a different model. Distilling a model you own is routine engineering; the adversarial variant harvests a competitor's protected traces through its public interface.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

GPT-6 Astra Goes to Work: OpenAI's Priciest Model Bets Everything on Computer Use
LLM & Chatbots

GPT-6 Astra Goes to Work: OpenAI's Priciest Model Bets Everything on Computer Use

OpenAI has begun rolling out GPT-6 Astra to business customers, framing its newest frontier model less as a chatbot and more as a worker that operates software...

Seung Jung21 days ago
ChatGPT Images 2.5 Halves Generation Latency and Adds a Sketch Canvas
LLM & Chatbots

ChatGPT Images 2.5 Halves Generation Latency and Adds a Sketch Canvas

OpenAI shipped ChatGPT Images 2.5 with up to 50% lower latency, a Sketch drawing canvas, templates, pinned comments and two new API models.

Seung Jung20 days ago
Sol and Luna Halve OpenAI's Mid-Tier Token Prices, But GPT-5.6 Still Wins Two Charts
Developer Tools

Sol and Luna Halve OpenAI's Mid-Tier Token Prices, But GPT-5.6 Still Wins Two Charts

OpenAI shipped two mid-tier models on Tuesday and made the pitch almost entirely about money. GPT-6 Sol lists at $2 per million input tokens and $10 per million...

Seung Jung8 days ago
OpenAI's Astra for Law Is Not a New Model. It Is a 230-Million-URL Search Index
LLM & Chatbots

OpenAI's Astra for Law Is Not a New Model. It Is a 230-Million-URL Search Index

OpenAI's Astra for Law wraps GPT-6 Astra in a 230-million-URL legal index, lifting Vals AI Legal Research Bench correctness from 38.7% to 54%.

Seung Jung13 days ago
The Best Model on OpenAI's New Mental Health Benchmark Scores 57.3%
LLM & Chatbots

The Best Model on OpenAI's New Mental Health Benchmark Scores 57.3%

OpenAI's MentalHealthBench, built with 80-plus clinicians in 22 countries, tops out at 57.3% for GPT-6 Astra. Claude Opus 5.5 scored 52.4%.

Seung Jung3 days ago
OpenAI Ships a Data Agent and a Wall Street ChatGPT on the Same Day
LLM & Chatbots

OpenAI Ships a Data Agent and a Wall Street ChatGPT on the Same Day

OpenAI launched a Data agent for ChatGPT Work and ChatGPT for Financial Services, built with Morgan Stanley and Evercore, on the same Thursday.

Seung Jung20 days ago