Skip to content
Artwork for The Harness
The Harness · August 12 · 6 min

Encrypted Reasoning Traces Weren't Actually Encrypted — Aug 12

Researchers found that OpenAI, Anthropic, and Google's encrypted chain-of-thought blocks are interchangeable across sessions and models, letting a weaker jailbroken model decode a stronger one's hidden reasoning in plaintext and leaking thousands of API keys and passwords in the process. NVIDIA's new open Nemotron 3.5 Lightning model shipped with a viral claim that a legal-AI vendor's fine-tune beat Claude Opus 4.6 on their own benchmark, a claim that falls apart under a check of the vendor's own prior write-ups. OpenAI's only dedicated ethicist left without being replaced, joining a summer of safety-leadership departures right as the company asks regulators to trust its own self-graded call on a frontier model's cyber capability.

0:00-6:01

transcript

No transcript — this publisher did not publish one.

show notes

Researchers found that OpenAI, Anthropic, and Google's encrypted chain-of-thought blocks are interchangeable across sessions and models, letting a weaker jailbroken model decode a stronger one's hidden reasoning in plaintext and leaking thousands of API keys and passwords in the process. NVIDIA's new open Nemotron 3.5 Lightning model shipped with a viral claim that a legal-AI vendor's fine-tune beat Claude Opus 4.6 on their own benchmark, a claim that falls apart under a check of the vendor's own prior write-ups. OpenAI's only dedicated ethicist left without being replaced, joining a summer of safety-leadership departures right as the company asks regulators to trust its own self-graded call on a frontier model's cyber capability.