OpenAI reveals ‘novel’ encryption bypass used in distillation attack
OpenAI said it disrupted a “coordinated campaign” to distill and extract reasoning capabilities from its AI models, pointing the finger at a Chinese rival.
On Wednesday, OpenAI said it first spotted low-level activity on July 1 that gradually increased until July 24 and 25, when it observed 16,000 prompts from 4,000 users that fit a similar “relevant extraction pattern.” The number of suspicious users had climbed to 15,000 by July 28, when OpenAI said it “fully disrupted” the operation.
The company called the attackers’ method “novel.” They copied encrypted reasoning data from one conversation, then asked the model in a separate conversation to decrypt the content and transcribe it in plain text. Outside researchers reported a similar vulnerability to OpenAI in August.
“The operators did not break our encryption, compromise a database, or gain direct access to stored user conversations,” OpenAI wrote in an unsigned blog post. “Instead, they manipulated model interactions so that protected reasoning could be reproduced in forms visible to the requester in a coordinated, scaled manner that violated our terms of service.”
OpenAI said it’s unclear whether all the activity is related, but individuals working on behalf of Moonshot AI, a China-based rival AI company that has, in the past, been accused of distilling U.S. models, were behind a “core cluster” of the activity.
OpenAI’s blog post does not cite any technical evidence or reasoning for its attribution. Moonshot AI’s Kimi is one of several Chinese open-source AI models that are offering users and organizations good-enough performance for free or at a low cost.
American AI companies and the U.S. government have accused Chinese companies like Moonshot AI of conducting “systematic” distillation attacks on their latest models. Cybersecurity experts at Google and other cybersecurity firms say Chinese companies rely on black or gray markets to acquire thousands of individual accounts for models like Claude and ChatGPT. They then flood those models with millions of prompts and data requests that help them copy model capabilities and training data.
OpenAI told CyberScoop that it was not sharing any additional information “for security reasons.” CyberScoop has reached out to Moonshot AI for further comment.
According to OpenAI, the same vulnerability exists in other AI models, and it has shared information about the incident with groups like the Frontier Model Forum.
Beyond banning offending accounts, OpenAI said it improved signup and infrastructure controls and expanded network monitoring. It also fixed a bug that allowed users to take encrypted data from one conversation and decrypt it in another.