Stealing Reasoning Traces from Proprietary LLM APIs: A Critical Vulnerability Fixed
Discovery of a Critical Vulnerability in Reasoning Traces
FAQ
What are encrypted reasoning traces?
They are encrypted data blocks returned by models like GPT-5.6 and Claude, containing internal reasoning steps, sent to clients for transparency without exposing content.
How was the vulnerability exploited?
Researchers replayed encrypted blocks on weaker models from the same family (e.g., Claude Haiku 4.5) and used prompt injection to jailbreak them and extract raw reasoning.
Has the vulnerability been fixed?
Yes, all providers (OpenAI, Anthropic, Google) acknowledged the report, and researchers could not repeat the attack after the fix.
What impact does this have on MENA enterprises?
Organizations relying on proprietary models should reassess security and privacy policies, ensuring systems are updated with the latest patches.
Source: Simon Willison (LLM & tools)
AI-assisted content, human-reviewed.