AI/News

OpenAI says a Moonshot-linked cluster sent 16,000 requests in two days to pull its models' hidden reasoning; 15,000 accounts cut off by 28 July

The campaign ran from 1 July. OpenAI says no encryption was broken and its figures count attempts, not successes; it attributes a core cluster to 'individuals associated with Moonshot AI', the Kimi developer, not to the company.

By Daily Aletheia · Checked against the primary source · 1 October 2026 · 2 min read


Image: OpenAI - announcement card for 'Disrupting a coordinated model-distillation campaign', openai.com, 30 September 2026

OpenAI published "Disrupting a coordinated model-distillation campaign" on 30 September.

What it says happened

OpenAI says it identified and disrupted a coordinated campaign to extract protected reasoning from its models - the model's internal working on a task, which is withheld from the final answer. It calls the activity consistent with adversarial distillation: the unauthorised use of one model's outputs or reasoning to train, reproduce or improve another.

The operators did not break encryption, compromise a database or reach stored user conversations. They manipulated model interactions so that protected reasoning was reproduced in forms visible to the requester, including by copying encrypted reasoning from one conversation and asking a model in another conversation to decrypt and transcribe it.

Activity began on 1 July at low volume. On 24 and 25 July OpenAI saw spikes of 16,000 requests using the extraction pattern from over 4,000 users. A wider cluster of more than 15,000 users showed related prompt patterns; OpenAI says it fully disrupted it by 28 July. A footnote states the figures describe attempted, not necessarily successful, extractions.

Independent security researchers separately disclosed related cross-model and conversation-compaction vulnerabilities, which OpenAI says it confirmed were real.

Attribution

OpenAI says it is unclear whether all operators came from a single actor, but attributes "a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi".

What it did

Accounts were banned or restricted, signup and infrastructure controls strengthened, and monitoring expanded. OpenAI closed a pathway that let someone holding another user's encrypted reasoning replay it and recover the contents, and added checks to hold streamed output that might expose reasoning. It shared findings through the Frontier Model Forum and government information-sharing channels, and says the manipulation is not unique to its models.

OpenAI's stated concern is that extracted reasoning could train another model without the safeguards applied to the original's user-facing outputs, and that distillation at scale transfers capabilities without the same investment in safety. It says the work is not finished: partner-hosted deployments need the same protections, and tool-output attacks need protections beyond ordinary visible text.

The Check

What was said
"OpenAI claims Moonshot extracted its data to train AI models" (The Verge, 30 September 2026)
What the Source Says
OpenAI attributes "a core cluster of the activity to individuals associated with Moonshot AI"; its footnote says the figures "describe attempted, not necessarily successful, extractions"; the post calls the activity "consistent with adversarial distillation" and does not say what, if anything, was trained on it.
The Difference
The page names individuals linked to Moonshot for part of the activity and does not claim the reasoning was used to train a model. "Moonshot extracted its data to train AI models" states as done, and as the company, what OpenAI states as attempted and as individuals.

Sources & further reading

  1. 01OpenAI - Disrupting a coordinated model-distillation campaign, 30 Sept 2026 ↗Primary
  2. 02The Verge - 'OpenAI claims Moonshot extracted its data to train AI models', 30 Sept 2026 (the claim checked) ↗Attributed