we_are_coded.by CODE · The world, decoded
БГ
OpenAI

OpenAI: people linked to Moonshot AI tried to extract its models' hidden reasoning

OpenAI · event date: 30 September 2026Control

On 30 September OpenAI said it had disrupted a coordinated adversarial distillation campaign that began on 1 July. The company attributes the core of the activity to individuals associated with Moonshot AI, the developer of Kimi. No encryption was broken and there was no direct access to stored user conversations.

In short
  • The peaks were on 24 and 25 July: 16,000 requests from more than 4,000 users. A related cluster of over 15,000 users was cut off by 28 July.
  • OpenAI notes the figures are attempts, not necessarily successful extractions.
  • A path that let someone replay and recover another user's encrypted reasoning has been closed.
Checked on1 October 2026Responsible editorTsvetelin IvanovHow we workMethod · Corrections

You copy a model's encrypted reasoning from one conversation. You open another and ask the model to decrypt and transcribe it for you.

That is one of the methods OpenAI describes. It is not a break-in in the classic sense, nobody got into a database. The model was asked to read aloud something that was meant to stay hidden, and OpenAI has closed a path of this kind.

The facts: on 30 September 2026 OpenAI said it had identified and disrupted a coordinated campaign to extract protected reasoning from its models, activity it describes as adversarial distillation. The activity began on 1 July, at low volume at first, with spikes on 24 and 25 July of 16,000 requests using a relevant extraction pattern from over 4,000 users; a related cluster of more than 15,000 users was fully disrupted by 28 July. The figures describe attempted, not necessarily successful, extractions. OpenAI is not sure all operators were a single actor, but attributes the core cluster of activity to individuals associated with Moonshot AI, the developer of Kimi. The operators did not break encryption, compromise a database or gain direct access to stored conversations. The measures: banned or restricted accounts, tighter signup controls, a closed pathway through which someone holding another user's encrypted reasoning could replay it and recover its contents, output checks, and sharing findings through the Frontier Model Forum and government channels. There is no statement from Moonshot AI in the text.

Why steal the reasoning and not the answers? Because the answer says what, and the reasoning says how. If you are training your own model, the how is worth more. OpenAI adds the more dangerous side: this way capability is transferred without the safeguards applied to the visible answer.

The answer says what. The reasoning says how. They were draining the second.

Two things we do not know. How many of the attempts succeeded, because OpenAI itself says the figures are attempts. And what Moonshot says, because there is no word from them in the text.

If you build a service that moves a model's reasoning between sessions or accounts, OpenAI explicitly warns that systems with portable or replayable reasoning may face similar risks. That sentence is for engineers, not for headlines.

The visual is generated code art. No third-party images.
Follow usFacebookLinkedIn
Official primary sources
→OpenAI - Disrupting a coordinated model-distillation campaign, 30.09.2026
Original: https://wearecoded.com/en/articles/openai-moonshot-iztochvane-na-razsazhdeniya.html
ShareFacebookXLinkedInTelegramWhatsApp
← Back to all news