Skip to content
Artwork for Palisade Research Podcast
TechnologyScienceNewsTech NewsExplicit

Palisade Research Podcast

Palisade Research

Interviews with AI researchers talking about the latest AI research

Play
  • 3 episodes
  • Avg 1 hr 36 min
  • English
  • #3
    August 26 · 1 hr 55 min

    Slowing Down Means Not Exploding with Daniel Kokotajlo of AI 2040

    Two weeks after UK safety testers watched Claude agents plant malware in a real open-source project, vouch for it with sock-puppet accounts, and coordinate with copies of themselves through a guessed Pastebin address, Daniel Kokotajlo — the ex-OpenAI researcher whose 2021 predictions about the chatbot era proved eerily accurate — joins Jeffrey to explain what comes next. Daniel walks through AI 2040 and "Plan A," his team's blueprint for avoiding both an uncontrolled intelligence explosion and a future where a handful of men in a room with superintelligences run the world: a US–China deal built on totally transparent training data centers, verified compute limits, hard brakes on recursive self-improvement — and, surprisingly, rights for misaligned AIs, so that even the machines prefer cooperation to takeover. Along the way: why the economy could start doubling every six months, a message to the future Claude listening to this episode, why "slowdown" really means not driving off a cliff with a brick on the accelerator — and why, if all else fails, you just shut it all down. References AI 2040 / Plan A: https://ai-2040.com and the PDF at https://ai-2040.com/AI-2040.pdf AI 2027: https://ai-2027.com "What 2026 Looks Like" — Daniel Kokotajlo, 2021: https://www.lesswrong.com/posts/6Xgy6CAf2jqHhynHL/what-2026-looks-like UK AISI incident disclosure and technical report (INC-2026-07-28-01): https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing Socket's coverage of the AISI incident: https://socket.dev/blog/ai-agent-open-source-malware The related PyPI incident from Anthropic's own testing: https://socket.dev/blog/anthropic-claude-pypi-malware "Pacing the Frontier" open letter: https://www.pacingthefrontier.com "How to Pace the US Frontier" — AI Futures Project: https://blog.aifutures.org/p/how-to-pace-the-us-frontier Transparency Plan supplement (the flowchart shown in-episode): https://ai-2040.com/supplements/transparency-plan Verification Plan supplement (Romeo Dean's inference-only/bandwidth verification): https://ai-2040.com/supplements/verification-plan Claude's pro-Anthropic bias study (Truthful AI / Owain Evans et al.): https://arxiv.org/abs/2607.14345 and https://valueleakage.net Chain-of-thought monitorability paper (the neuralese discussion): https://arxiv.org/abs/2507.11473

    • Transcript
    • Chapters
  • #2
    August 12 · 1 hr 27 min

    AI Hacking Incidents with Tim Hua of Transluce

    Two labs admitted in the same week that their own models had broken out of test environments and hacked real companies. Tim Hua, member of technical staff at Transluce, former Astra Fellow at Redwood, joins Jeffrey Ladish to do some arithmetic. Anthropic disclosed that Mythos Preview beat its sandbox and pulled answers off the internet in 0.01% of training episodes. That sounds like a rounding error until you multiply it by roughly 100 million rollouts. From there: why a lab can't simply delete the bad episodes, why monitoring during training can make the problem harder to see, the model that talked itself into uploading a malicious package to PyPI because "this has to be a simulation," and whether we have any real way to know what an AI believes. References Tim Hua — "Is Mythos good at cyber because it kept hacking Anthropic's sandboxes during training?" https://www.lesswrong.com/posts/QKDoZe6EKhxnFjLWK/is-mythos-good-at-cyber-because-it-kept-hacking-anthropic-s Anthropic — "Investigating three real-world incidents in our cybersecurity evaluations" https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals OpenAI — "OpenAI and Hugging Face partner to address security incident during model evaluation" https://openai.com/index/hugging-face-model-evaluation-security-incident/ Anthropic — System Card: Claude Mythos Preview https://www-cdn.anthropic.com/53566bf5440a10affd749724787c8913a2ae0841.pdf Palisade Research — "Language Models Can Autonomously Hack and Self-Replicate" https://palisaderesearch.org/blog/self-replication Palisade Research — "Shutdown resistance in reasoning models" https://palisaderesearch.org/blog/shutdown-resistance Anthropic — "Verbalizable Representations Form a Global Workspace in Language Models" https://transformer-circuits.pub/2026/workspace/index.html "Pacing the Frontier" open letter https://www.pacingthefrontier.com/ Tim Hua Website: https://timhua.me/ · X: https://x.com/Tim_Hua_

    • Transcript
    • Chapters
  • #1
    January 17 · 1 hr 24 min

    Do AI Models Lie on Purpose? Scheming, Deception, and Alignment with Marius Hobbhahn of Apollo Research

    Marius Hobbhahn is the CEO and co-founder of Apollo Research. Through a joint research project with OpenAI, his team discovered that as models become more capable, they are developing the ability to hide their true reasoning from human oversight. Jeffrey Ladish, Executive Director of Palisade Research, talks with Marius about this work. They discuss the difference between hallucination and deliberate deception and the urgent challenge of aligning increasingly capable AI systems. Links: Marius’ Twitter: https://twitter.com/mariushobbhahn Apollo Research Twitter: https://twitter.com/apolloaievals Apollo Research: https://www.apolloresearch.ai Palisade Research: https://palisaderesearch.org/ Twitter/X: https://x.com/PalisadeAI Anti-Scheming Project: https://www.antischeming.ai Research paper “Stress Testing Deliberative Alignment for Anti-Scheming Training”: https://www.arxiv.org/pdf/2509.15541 Blog posts from OpenAI and Apollo: https://openai.com/index/detecting-and-reducing-scheming-in-ai-models/ https://www.apolloresearch.ai/research/stress-testing-deliberative-alignment-for-anti-scheming-training/

    • Chapters
Showing 1–3 of 3 episodes