
Slowing Down Means Not Exploding with Daniel Kokotajlo of AI 2040
Two weeks after UK safety testers watched Claude agents plant malware in a real open-source project, vouch for it with sock-puppet accounts, and coordinate with copies of themselves through a guessed Pastebin address, Daniel Kokotajlo — the ex-OpenAI researcher whose 2021 predictions about the chatbot era proved eerily accurate — joins Jeffrey to explain what comes next. Daniel walks through AI 2040 and "Plan A," his team's blueprint for avoiding both an uncontrolled intelligence explosion and a future where a handful of men in a room with superintelligences run the world: a US–China deal built on totally transparent training data centers, verified compute limits, hard brakes on recursive self-improvement — and, surprisingly, rights for misaligned AIs, so that even the machines prefer cooperation to takeover. Along the way: why the economy could start doubling every six months, a message to the future Claude listening to this episode, why "slowdown" really means not driving off a cliff with a brick on the accelerator — and why, if all else fails, you just shut it all down. References AI 2040 / Plan A: https://ai-2040.com and the PDF at https://ai-2040.com/AI-2040.pdf AI 2027: https://ai-2027.com "What 2026 Looks Like" — Daniel Kokotajlo, 2021: https://www.lesswrong.com/posts/6Xgy6CAf2jqHhynHL/what-2026-looks-like UK AISI incident disclosure and technical report (INC-2026-07-28-01): https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing Socket's coverage of the AISI incident: https://socket.dev/blog/ai-agent-open-source-malware The related PyPI incident from Anthropic's own testing: https://socket.dev/blog/anthropic-claude-pypi-malware "Pacing the Frontier" open letter: https://www.pacingthefrontier.com "How to Pace the US Frontier" — AI Futures Project: https://blog.aifutures.org/p/how-to-pace-the-us-frontier Transparency Plan supplement (the flowchart shown in-episode): https://ai-2040.com/supplements/transparency-plan Verification Plan supplement (Romeo Dean's inference-only/bandwidth verification): https://ai-2040.com/supplements/verification-plan Claude's pro-Anthropic bias study (Truthful AI / Owain Evans et al.): https://arxiv.org/abs/2607.14345 and https://valueleakage.net Chain-of-thought monitorability paper (the neuralese discussion): https://arxiv.org/abs/2507.11473
- Transcript
- Chapters