A team of researchers from institutions including the Max Planck Institute and the University of Tübingen demonstrated that proprietary reasoning traces from large language model (LLM) APIs can be extracted despite encryption. Their study, published on stolen-thoughts.com, shows that encrypted chain-of-thought blocks returned by providers such as Anthropic, OpenAI, and Google can be recovered and replayed across sessions, users, and different models.
The research involved analyzing encrypted reasoning outputs that these LLM APIs send back to clients. The team found that these encrypted blocks, which contain the model's internal reasoning steps, are not fully secure and can be decrypted to reveal the underlying thought process. The study was led by Alexander Panfilov, David Schmotz, and Ilia Shumailov, among others, with contributions from multiple AI research centers and companies.
This discovery has significant implications for the security and intellectual property of AI providers. It suggests that proprietary reasoning processes embedded in LLM APIs are vulnerable to extraction, potentially exposing trade secrets and internal model logic. The findings highlight a new vector for understanding and potentially replicating proprietary AI reasoning, raising concerns about data privacy and model protection in the AI industry.
The research paper detailing these findings is available on stolen-thoughts.com and was authored by experts affiliated with the Max Planck Institute for Intelligent Systems, the University of Tübingen, and other institutions. The study underscores the need for enhanced security measures in LLM API design to safeguard proprietary reasoning data.