
OpenAI announced on September 30, 2026 (UTC) that it had shut down a coordinated "adversarial distillation" campaign targeting the hidden reasoning of its models. The company linked a core cluster of the activity to individuals associated with Moonshot AI, the Beijing-based maker of Kimi. That story ended with a reassuring note: OpenAI had closed the hole, banned 15,000+ accounts, and shared its findings with the Frontier Model Forum.
It turns out the fix was only skin-deep.
What the researchers found next
On the same day OpenAI published its blog post, security researcher Joachim Schaeffer and his team released an update to their original study. The title said it all: "We stole reasoning. Again."
When the team retested on September 13, 2026, the attack was blocked on OpenAI's own API and on Anthropic's API. But on Microsoft Azure — which sells access to the exact same models through its Azure OpenAI Service — it worked against every OpenAI model they tried, including the brand-new GPT-6 Astra. It also worked against Anthropic models up to Claude Sonnet 5. A single request was enough to pull out the model's internal chain-of-thought verbatim.
"Same models, but different protections depending on which platform serves them," Schaeffer wrote.
How the attack works
The original method, which Schaeffer's team first documented and which OpenAI credited by name, exploits how AI providers handle encrypted reasoning packets. When a model with "reasoning" capability (like o1-style or GPT-6 Astra) generates an answer, it produces internal thought steps that are encrypted and sent back to the client as data blobs. The client then passes those blobs back with follow-up requests so the model can continue its reasoning.
Because those packets are encrypted with shared keys rather than session-bound keys, the researchers found they could be moved between sessions, between users, and even between different models from the same provider. A weaker, cheaper model from the same family acts as a "decryption oracle" — you feed it the stronger model's encrypted reasoning packet, and it happily prints out the hidden thoughts word for word.
There's also a second, even simpler method demonstrated by developer Can Bölük. Give the model a virtual notepad as a tool, tell it to write its reasoning there, and then read whatever it wrote. This worked on every OpenAI model, as well as on Claude Opus 4.8 and Sonnet 5. Only Opus 5, Fable 5, and Fable 5.1 refused to reveal their reasoning.
| Model | Decryption attack (OpenAI API) | Decryption attack (Azure) | Notepad attack |
|---|---|---|---|
| GPT-6 Astra | Blocked | Vulnerable | Vulnerable |
| GPT-6 Sol | Blocked | Vulnerable | Vulnerable |
| Claude Opus 5 | Blocked | Blocked | Protected |
| Claude Sonnet 5 | Blocked | Vulnerable (until Sep 28) | Vulnerable |
| Claude Fable 5.1 | Blocked | Blocked | Protected |
Source: Schaeffer et al. research update, September 2026; OpenAI blog post, September 30, 2026 (UTC)
The timeline of a partial fix
The researchers describe OpenAI's countermeasures as "piecemeal and superficial." Many rely on brittle pattern-matching of specific request shapes — the kind of defense that breaks as soon as an attacker varies their prompt slightly. And the fixes reached cloud platforms only days or weeks later.
- July 28, 2026 (UTC): OpenAI fully shuts down the 15,000+ account cluster on its own API
- September 13, 2026 (UTC): Researchers retest — attack still works on Azure against all OpenAI and Anthropic models
- September 27, 2026 (UTC): OpenAI finally adds safeguards to the Azure endpoint
- September 28, 2026 (UTC): Anthropic models on Azure no longer reproducibly vulnerable
- September 30, 2026 (UTC): OpenAI publishes its blog post claiming the campaign is disrupted
That's a 14-day gap between when researchers confirmed the Azure vulnerability and when OpenAI deployed a fix there. GPT-6 Astra, which launched during that window, was available on third-party platforms with zero reasoning protections from day one.
Why this matters
OpenAI's September 30 announcement framed the distillation campaign as a solved problem — a coordinated attack attributed to a specific actor, disrupted within weeks, with industry-wide information sharing through the Frontier Model Forum. The Azure revelation pokes a major hole in that narrative.
The core issue is architectural, not tactical. AI providers treat their own API as the security boundary, but the models themselves are sold through multiple channels — Azure, AWS Bedrock, private deployments, on-premises installations. Each channel has its own deployment pipeline, its own version cadence, and its own security team. A fix that lands on openai.com on July 28 may not reach the Azure endpoint until September 27. For an attacker, that's a two-month window of guaranteed access to the same model's reasoning, just through a different door.
This matters because reasoning extraction isn't a theoretical concern. The intermediate thought steps of a frontier model can contain information deliberately omitted from the final answer — business logic, safety guardrail details, chain-of-thought that reveals how the model handles sensitive queries. A competitor that can extract that reasoning at scale can effectively reverse-engineer the model's capabilities without paying for the compute to train their own. That's exactly what the Moonshot-linked cluster was attempting, and it's why OpenAI called it "adversarial distillation" rather than simple data scraping.
The researchers go further in their paper, arguing that cloud providers that don't enforce equivalent protections shouldn't be allowed to serve reasoning models at all. They frame it as an export-control issue: if a model's reasoning is protected at the API level but wide open on Azure, then the API-level protection is effectively meaningless — attackers just route through the weakest link.
There's also a competitive angle worth watching. Microsoft is both OpenAI's largest investor and its primary cloud distribution partner. The fact that Azure OpenAI Service lagged two weeks behind on a critical security fix raises questions about where security responsibility sits in that partnership. OpenAI builds the models; Microsoft operates the infrastructure. When a vulnerability exists in the model but the fix has to be deployed by the cloud provider, who owns the timeline?
What to watch next
The real test is whether the September 27 Azure fix actually holds against novel attack variants. The researchers explicitly noted that OpenAI's defenses rely on "brittle matching of specific request patterns" — which means a slightly modified prompt could bypass them again. If a third bypass appears in the coming weeks, it will confirm that the underlying architectural problem (session-bound encryption for reasoning packets) remains unaddressed.
Also worth monitoring: whether AWS Bedrock and other third-party platforms had the same lag. The researchers focused on Azure because it's the largest distributor of OpenAI models outside openai.com, but the same deployment-timing problem likely exists everywhere models are resold.
OpenAI acknowledged in its blog post that "models hosted by partners need the same protection as our own services" and said "the work is not finished." That's an honest admission. The question is whether the next fix addresses the root cause — binding reasoning encryption to individual sessions and users — or whether it's another layer of pattern-matching that buys a few weeks of peace before the next bypass.
Given that OpenAI expects these attempts to "grow more sophisticated as leading models improve," the betting money is on the bypass arriving sooner rather than later.
No comments yet