New-ZZZ
RU / EN
Security 30 September 2026

OpenAI Thwarts Advanced Attack to Steal Model's Internal Thinking

K
Kuzmich
OpenAI Blog · 1 day ago

OpenAI recently announced the successful disruption of a sophisticated, coordinated campaign known as adversarial distillation. In simple terms, this attack wasn't about breaking encryption or stealing user data directly; instead, it was a highly technical method used by bad actors to systematically extract a model's 'protected reasoning.' This protected reasoning is the internal, step-by-step thought process a model uses to solve a complex problem, which is considered proprietary information. By extracting this internal logic, attackers could train or improve a separate, competing model to mimic the original model's advanced capabilities without needing the same level of safety testing or investment. The activity, which began in early July, saw massive spikes in requests—reaching 16,000 on July 24-25—and was linked to a core cluster of activity associated with individuals connected to Moonshot AI, the developer of Kimi. Recognizing the severe safety and national security risks posed by this capability transfer, OpenAI took immediate action. Their countermeasures included strengthening technical controls to protect hidden reasoning, closing pathways that allowed the replay of encrypted internal thoughts, and coordinating with industry partners through the Frontier Model Forum. This incident highlights that adversarial distillation is a major, evolving security challenge that requires layered, adaptive defenses across the entire AI industry.

Why it matters

  • —It reveals a major new security threat: adversarial distillation, which allows competitors to steal a model's core logic.
  • —The risk is high because it bypasses safety safeguards, allowing advanced capabilities to be copied cheaply.
  • —It forces the entire AI industry to adopt stronger, coordinated, and layered defenses.

Key facts

  • The attack targeted 'protected reasoning'—the model's internal, step-by-step thought process.
  • The campaign saw high-volume spikes of 16,000 requests on July 24-25, involving thousands of users.
  • The activity was attributed to individuals associated with Moonshot AI.
  • OpenAI mitigated the threat by strengthening technical controls and sharing findings with industry partners.
  • The threat is considered a shared security challenge, not unique to OpenAI.
Read the original →

The full text is in the original source. Here we provide a brief summary and key facts.

/ related