

What's Inside Anthropic's 186-Page Risk Report
For close to a year, the bioweapon filters at one of the world's strongest AI companies had a blind spot. One internal flag had switched off both the blocking and the logging for an entire pipeline — roughly 50,000 hired contractors, 133 million exchanges. No regulator caught it. No law required anyone to tell you. You know because Anthropic wrote it down, inside a 186-page risk report released with no announcement, no author's name, and a title that literally says Redacted. This episode reads the whole document — and asks the question it leaves behind: can a confession be a system? What's inside: the re-scan of a year of unwatched conversations — 1,197 flagged, 62 real after the company's own red-team traffic is removed, zero misuse found — and why that answer is careful, not clean. The frontier model Anthropic finished and won't release, stronger than its flagship, working only inside the building — and the 2019 GPT-2 withholding playbook it repeats, written in part by the people who later founded Anthropic. The strangest audit ever run: Anthropic handed the misalignment chapter to Claude itself, which led with its own conflict of interest, called the report candid and largely faithful, and got two criticisms into the final text. The case against gets a full act: passing thresholds rewritten mid-stream, an oversight trust whose outside-review power has never been used — and whose trustee left to become Anthropic's Chief Global Affairs Officer — and the analyst who read all 186 pages and wasn't reassured. Then the machinery question: aviation buys pilots' honesty with immunity and a ten-day clock; every AI reporting law in force routes truth to agencies confidentially. Three dated calls close it out — one already landed. CHAPTERS 00:00 The year the filters were off 01:15 Disclosure: this show runs on the subject's AI 02:16 A 186-page confession nobody signed 03:01 One flag — the brake and the camera 04:21 The re-scan: 1,197 flags, 62 real 05:38 Anthropic re-grades itself 06:15 The model they won't release 07:16 Why hold a frontier model back 09:10 The 2019 playbook, repeated 10:22 The strangest audit ever run 11:52 Seven more confessions 13:04 The case against candor 15:14 The machinery that buys honesty 17:26 Three holdbacks in a fortnight 17:58 Capture — and the trustee who switched sides 19:31 The ledger: three dated calls 20:22 Nobody had to tell you SOURCES Anthropic — Redacted Risk Report, August 2026 (PDF): https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf Anthropic — RSP updates: https://www.anthropic.com/rsp-updates Zvi Mowshowitz — On the Anthropic Risk Report: https://thezvi.substack.com/p/anthropic-risk-report-august-2026 OpenAI — GPT-2 Release Strategies (2019): https://arxiv.org/abs/1908.09203 Reuters — Anthropic's letter alleging Alibaba extraction: https://www.reuters.com/world/china/anthropic-says-alibaba-illicitly-extracted-claude-ai-model-capabilities-2026-06-24/ TechCrunch — OpenAI slows Astra over security concerns: https://techcrunch.com/2026/08/07/openai-says-it-slowed-astra-model-development-over-security-concerns/ Unite.AI — Z.ai's GLM-5.3 cyber capability: https://www.unite.ai/z-ai-launches-glm-5-3-with-frontier-coding-and-a-cyber-capability-that-outgrew-its-training/ Anthropic — Tino Cuéllar joins: https://www.anthropic.com/news/tino-cuellar NASA ASRS — immunity provisions: https://asrs.arc.nasa.gov/overview/immunity.html NY RAISE Act (S6953B): https://www.nysenate.gov/legislation/bills/2025/S6953/amendment/B EU AI Act, Article 55: https://artificialintelligenceact.eu/article/55/


















