Skip to content
Artwork for Turing Post

Turing Post

Turing Post

Hi, I’m Ksenia, founder of Turing Post.


On this channel, I talk to the people shaping AI and pay attention to the ideas, shifts, and details others might miss.


Inference is my interview show with innovators, builders, founders, and thinkers moving AI forward.


Attention Span is where I slow down on what deserves a closer look: the signals, questions, and stories hiding between the headlines.


Subscribe for the unusual takes. And always stay curious!

Play
  • 20 episodes
  • Avg 19 min
  • English
  • September 14 · 17 min

    AI That Acts: Devin Fusion, Persimmon, Programmable Worlds & Amodei’s Slowdown Call

    Can AI-generated worlds remember what happened off-screen? How do dual-agent setups reduce costs? In this episode of Attention Span, we break down the biggest shifts in AI’s ability to understand, reason, and act in the physical and digital world. We dive into Alaya’s programmable world model, real-world vs. simulated robotics with Telexistence and Skild S1, Cognition’s Devin Fusion and SWE-2, and the heated frontier slowdown debate between Dario Amodei and David Sacks. Hosted by Ksenia Se (founder of Turing Post). Covering September 7–13, 2026. 👇 Which topic should we cover in a deep-dive episode? Let us know in the comments! Also, this is the first episode of Attention Span Weekly News – let us now if you like us to continue doing this. 🔳 SPONSORSHIP Partner with Turing Post to reach top AI engineers, researchers, and builders: ks@turingpost.com 🔷 WORLD MODELS Alaya PWM paper: https://arxiv.org/abs/2609.10540 Demonstrations: https://alaya-lab.github.io/pwm/ World Labs - Atlas: https://www.worldlabs.ai/blog/atlas World in World: https://arxiv.org/abs/2609.11548 🔶 ROBOTS & SIMULATION AWS and Telexistence - DreamZero experiments: https://aws.amazon.com/blogs/physical-ai/bringing-a-frontier-world-model-to-the-convenience-store-inside-telexistences-dreamzero-experiment-on-aws/ NVIDIA - Skild S1: https://blogs.nvidia.com/blog/skild-ai-s1-physical-ai/ Skild’s S1 research: https://www.skild.ai/blogs/s1 Mila and Worldmodeldata: https://mila.quebec/en/news/worldmodeldata-and-mila-partner-to-prove-scaling-law-for-world-models 🔹 DEVIN FUSION & SWE-2 Cognition - Fusion: https://cognition.com/blog/local-fusion SWE-2: https://cognition.com/blog/swe-2 🔸 PERSIMMON & PERSONAL AGENTS Persimmon: https://persimmon.humansand.ai/blog/persimmon.html Grok Bot: https://x.ai/bot Meta Muse: https://about.fb.com/news/2026/09/introducing-muse-personal-ai-agent/ Muse security architecture: https://research.meta.ai/blog/security-and-safety-for-ai-agents-our-approach-with-muse Attention Span episode on Muse: https://www.youtube.com/watch?v=1pqzii7di7w 🔺 SAFETY, EVALUATORS & THE SLOWDOWN DEBATE Anthropic - cybersecurity alignment assessment: https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents Dario Amodei - We Must Pace the Frontier: https://darioamodei.com/post/we-must-pace-the-frontier Hugging Face Open Alignment Initiative: https://www.techmeme.com/260912/p13 David Sacks’s response: https://x.com/DavidSacks/status/2098973625252708460 🔻 MORE FROM TURING POST Newsletter, research coverage, and AI Builds AI: https://www.turingpost.com/ Instagram: https://www.instagram.com/turingpost_tv TikTok: https://www.tiktok.com/@turingpost_tv Subscribe for Monday news digests and Attention Span deep dives into how AI works. #WorldModels #AIAgents #TuringPost

  • September 13 · 16 min

    The Agent Can Rewrite Itself. So Who Controls It?

    Meta just gave its new personal agent, Muse, a remarkable amount of freedom. It can browse, work across your accounts, write code, build tools and run subagents. But Meta made one part of the system deliberately difficult for Muse to control: its own authority. In this episode of Attention Span, I look inside Muse Secure VM and the security architecture surrounding the agent. We get into Sentinel, the separate system that decides what Muse is allowed to do; why Muse can use your accounts without ever seeing the real credentials; how “tainted egress” changes permissions after a process touches private data; and why reading an email can give an agent considerably more power than “read access” suggests. And somehow, all of this takes us back to computer-security ideas from 1975. The larger question is one we are going to encounter everywhere as agents become more capable: How much freedom can we give an agent to discover new ways of doing things without allowing it to expand its own authority? *Watch it*, and tell me where you would draw that boundary. *IN THIS EPISODE* → What “rogue agent” actually means technically → How Muse Secure VM isolates the agent → Why Sentinel sits outside the environment Muse can modify → How an agent can write a new tool without granting that tool permission → How surrogate tokens keep real credentials away from the model → What “tainted egress” means and why permission may need memory → Why “read my email” can unlock much more than email → Least privilege and complete mediation, 51 years later → Where Muse's security architecture still has unresolved problems → Secure from whom? The agent, other users, or Meta itself → DeepSeek Harness + Muse: capability can become fluid; authority cannot *META MUSE* → Introducing Muse Meta's product announcement and overview of Muse Secure VM. https://about.fb.com/news/2026/09/introducing-muse-personal-ai-agent/ → How We Built Safety Into Muse The technical deep dive. This is where Meta explains Sentinel, Linux isolation, surrogate tokens, credential insertion, eBPF-based taint tracking and network egress controls. https://research.meta.ai/blog/security-and-safety-for-ai-agents-our-approach-with-muse *THE SECURITY IDEAS BEHIND IT* → Melanie Mitchell: Misleading Metaphors and Real Risks A useful grounding discussion of what we actually mean when we say an agent “escaped” or “went rogue.” https://aiguide.substack.com/p/misleading-metaphors-and-real-risks → Saltzer & Schroeder: The Protection of Information in Computer Systems The 1975 paper behind principles such as least privilege and complete mediation that suddenly look very current again. https://web.mit.edu/Saltzer/www/publications/protection/Basic.html *RELATED ATTENTION SPAN* → Everything Is a Plugin: DeepSeek Harness Our previous episode on agents that can create and modify their own tools. https://www.youtube.com/watch?v=jtyV7O4Pt0s → AI Escaped? The OpenAI–Hugging Face Incident What actually happened when cybersecurity agents found routes outside their intended environment. https://www.youtube.com/watch?v=RGeZ2moLkIc *MORE FROM TURING POST* → We Don't Know What They Know Our deeper look at the problem of understanding what increasingly capable models know and how they will use it. https://www.turingpost.com/p/we-don-t-know-what-they-know → Turing Post https://www.turingpost.com → Instagram @turingpost_tv https://www.instagram.com/turingpost_tv → TikTok @turingpost_tv https://www.tiktok.com/@turingpost_tv → Interviews: @realturingpost

  • September 9 · 33 min

    Does AI Understand the Machine It Runs On? | Inside OpenAI

    What if a model finds an optimization that surprises the engineers who have spent years working on that system? What, exactly, has it understood? I brought that question to OpenAI’s Phil Tillet and Matt Ferrari, whose work involves making AI cheaper and more accessible. They’re increasingly doing that work with the models themselves. Matt talks about research ideas his team used to dismiss because the engineering would be too complicated. Now they can give a model years of earlier research and ask it to explore what might work. They’re using models to help improve the smaller models behind speculative decoding, including the training process itself. So I ask whether the model is also helping choose the ideas, and how much that expands what they’re willing to try. I also ask something familiar to anyone who uses these systems: why does the same model sometimes feel different? Phil explains that even changing the order of floating-point calculations can introduce differences in its behavior. That puts a very concrete problem behind our conversation about understanding: an optimization can make the system faster and still change something you wanted to preserve. We get into how they catch those changes, and what happens when the failure is something nobody thought to test for. *We talk about:* When models became useful for engineering decisions. Why OpenAI’s models needed more control than Triton gave them. What happens inside the system after you send a request. How model-assisted kernel improvements helped cut Sol’s serving costs by 20%. Letting models investigate bugs independently, and deciding when to step in. Why successfully optimizing something can still be a waste of effort. Whether models need an internal representation of how a computer system behaves. How AI assistance opens up experiments that engineers previously couldn’t justify attempting. *Chapters:* *Follow on*: https://www.turingpost.com/ *Did you like the episode? You know the drill:* 📌 Subscribe here and here (https://www.turingpost.com/subscribe) for more conversations with the builders shaping real-world AI. 💬 Leave a comment 👍 Like it 🫶 Thank you for watching and sharing! *Guests:* Philippe (Phil) Tillet created Triton, a programming language that makes efficient GPU programming more accessible. He joined OpenAI as an intern in 2019, before it had an API or a product, and spent years improving training efficiency. His interests extend from compilers and kernels to Bertrand Russell and philosophy of mind. Matthew (Matt) Ferrari works on inference efficiency at OpenAI, across request routing, load balancing, debugging and speculative decoding. His fascination with optimization began in school, when GPU programming changed his understanding of how fast an algorithm could run. Today, he brings that curiosity to the entire system serving a model. #openai #inference #optimization

  • September 4 · 21 min

    NVIDIA’s $12.9B Plan to Rule Open-Source AI

    NVIDIA was once so hostile to open source that Linus Torvalds gave the company the finger. Today, it maintains open Linux modules, releases hundreds of models and datasets, and has reportedly agreed to buy Hugging Face for $12.9 billion. WHAT?! The change makes sense once we examine what NVIDIA learned from nearly dying with NV1, spending years searching for CUDA’s market, and watching researchers discover deep learning on gaming GPUs. In this episode, we follow that strategy from NV1 and CUDA to NVIDIA’s reported $12.9 billion acquisition of Hugging Face, and ask whether the company has found the most profitable model for open-source AI. *Watch it.* 👉 Subscribe for high-signal AI analysis 👉 Instagram https://www.instagram.com/turingpost_tv 👉 TikTok https://www.tiktok.com/@turingpost_tv 👉 More analysis: https://www.turingpost.com/ 👉 Interviews: @realturingpost Attention Span is here to explain the technical and business choices shaping AI. Links: NVIDIA FY2026 Form 10-K https://www.sec.gov/Archives/edgar/data/1045810/000104581026000021/nvda-20260125.htm Interview with Spencer Huang (Nvidia) https://www.youtube.com/watch?v=NEv9EnD7JVU&t=1231 Interview with Clem Delangue (Hugging Face) https://www.youtube.com/watch?v=DfJV722V1WY NVIDIA Open Source https://opensource.nvidia.com/en-us NVIDIA on Hugging Face https://huggingface.co/nvidia Reuters on the reported $12.9B agreement https://www.reuters.com/technology/nvidia-talks-acquire-hugging-face-13-billion-deal-business-insider-reports-2026-08-27/ NVIDIA Open GPU Kernel Modules https://github.com/NVIDIA/open-gpu-kernel-modules Turing Post’s history of computer vision and AlexNet https://www.turingpost.com/p/cvhistory6 #NVIDIA #OpenSourceAI #HuggingFace #AI #GitHub

  • August 31 · 16 min

    Fei-Fei Li, LeCun, Hassabis: What Do They Mean by “World Model”?

    Demis Hassabis, Yann LeCun, Fei-Fei Li – they all talk about “building a world model.” Some of them are dedicating their professional lives to it! But do they mean the same? So before joining the World Models workshop at Chicago Booth, I wanted to answer a basic question: what do researchers mean by a world model, and how many different ideas are sitting under this name? World models are absolutely fascinating area of research with its GPT moment still in the nearest future. This episode is based on the current research and provides a comprehensive overview of three broad approaches: generating future observations, predicting inside learned representations such as JEPA, and learning only what a planner needs to make decisions. *Watch it.* 👉 Subscribe for high-signal AI analysis 👉 Instagram https://www.instagram.com/turingpost_tv 👉 TikTok https://www.tiktok.com/@turingpost_tv 👉 More analysis: https://www.turingpost.com/ 👉 Interviews: @realturingpost Attention Span is here to show you AI isn’t magic. Sometimes the best way to understand a model is to change the background to purple and see what breaks. *Links:* Demis Hassabis on world models https://www.youtube.com/watch?v=sZaM6MadDZU Yann LeCun on world models https://www.youtube.com/watch?v=8sS9UJzb_t4 Fei-Fei Li on large world models https://www.youtube.com/watch?v=pNYVckbCFuk Beyond LLMs: JEPA and the Road to AGI – the main milestones so far https://www.youtube.com/watch?v=z0fh0SY3VWc stable-worldmodel https://github.com/galilai-group/stable-worldmodel/issues/153 VideoPhy-2, a benchmark https://arxiv.org/pdf/2503.06800 Physion-Eval https://arxiv.org/html/2603.19607v1 What Is JEPA? LeCun Architecture & World Models https://www.turingpost.com/p/jepa #WorldModels #AI #MachineLearning #YannLeCun #FeiFeiLi #DemisHassabis #JEPA #PhysicalAI #TuringPost #AttentionSpan

  • August 28 · 13 min

    OpenCode vs. OpenRouter: The Fight Over Your AI Models

    OpenCode began as an open-source coding agent. Now it is selling model access, negotiating directly with suppliers and preparing to reserve its own GPU capacity. That puts it on a collision course with OpenRouter, the model marketplace Stripe has agreed to acquire for a reported $8 billion. This episode follows this new shift in the industry and what Ox Alpha showed about the value of distribution: the company controlling the workflow may influence which models win long before a developer opens the model menu. *Watch it.* 👉 Subscribe for high-signal AI analysis 👉 Instagram https://www.instagram.com/turingpost_tv 👉 TikTok https://www.tiktok.com/@turingpost_tv 👉 Interviews: @realturingpost Attention Span is the video side of Turing Post. The newsletter goes to 115,000+ people who work on this stuff: https://www.turingpost.com #OpenCode #OpenRouter #AIAgents #CodingAgents #AIInfrastructure Sources and further reading OpenRouter is joining Stripe https://openrouter.ai/blog/announcements/openrouter-is-joining-stripe/ OpenCode https://opencode.ai/ Ox Alpha, Explained Without the Hype https://www.youtube.com/watch?v=tN8xiPoareo&t=16s OpenCode Zen https://opencode.ai/docs/zen/ GLM-5.3-Flash, formerly Ox Alpha, usage data https://opencode.ai/data/zhipuai/glm-5.3-flash Dax Raad on OpenCode’s direction https://x.com/thdxr/status/2093161006226612377 Dax Raad on inference economics https://x.com/thdxr/status/2093161006226612377 Dax Raad on OpenCode’s buying power https://x.com/thdxr/status/2092844520119345160 Jay V on OpenCode’s token volume https://x.com/snowmaker/status/2080667637861011924

  • August 25 · 17 min

    Ox Alpha, Explained Without the Hype

    An anonymous model called Ox Alpha appeared on OpenRouter and OpenCode on August 20 with a million-token context window, video input, and a price of zero. Within four days it had processed tens of trillions of tokens, and the internet had spent those same four days trying to work out who built it. In this episode: how you fingerprint a model you know nothing about, why the evidence points at Z.ai's unreleased multimodal GLM, what the 113-task benchmark runs really show versus the viral 80 percent, the three contradictory data policies governing your prompts, and the thought I keep coming back to – that the platform a model launches on is becoming as decisive as the lab that trained it. *Watch it.* 👉 Subscribe for high-signal AI analysis 👉 Instagram https://www.instagram.com/turingpost_tv 👉 TikTok https://www.tiktok.com/@turingpost_tv 👉 Interviews: @realturingpost Attention Span is the video side of Turing Post. The newsletter goes to 115,000+ people who work on this stuff: https://www.turingpost.com Sources and further reading Ox Alpha vs GLM-5.3 on OpenRouter: https://openrouter.ai/compare/stealth/ox-alpha/z-ai/glm-5.3 Ox Alpha on OpenCode https://opencode.ai/data/unknown/ox-alpha OpenCode Zen documentation https://dev.opencode.ai/docs/zen OpenRouter Stealth Model Terms https://openrouter.ai/terms/stealth The Tokenizer Is a Fingerprint by Joseph Elstner https://isimplifyme.com/whitepapers/the-tokenizer-is-a-fingerprint DeepSWE result https://x.com/winkey_h/status/2090814178810306874/photo/1 58.4% run, MatchaOnMuffins/oxalpha https://github.com/MatchaOnMuffins/oxalpha/blob/main/README.md 64.6% run, jyeric/ox-alpha-deepswe https://github.com/jyeric/ox-alpha-deepswe/blob/main/README.md Community fingerprinting summary: https://cellcog.ai/blog/what-is-ox-alpha/ Prediction market on the reveal: https://manifold.markets/Sketchy/who-is-behind-ox-alpha-the-mysterio #OxAlpha #OpenRouter #OpenCode #GLM #AIcoding #stealthmodel

  • August 25 · 15 min

    Etched Explained: The $21B AI Chip Startup Challenging NVIDIA

    Etched raised $1 billion in 26 days. Its valuation jumped from $10.3 billion to $21 billion. The second round was led by Jane Street after it tested Etched’s hardware and installed the first rack in its own data center. So what did Jane Street see? Etched began with Sohu, a Transformer-only ASIC that promised more than 500,000 tokens per second on Llama 70B. By 2026, Sohu and that claim had disappeared. Etched now sells a complete inference cluster and says it can run Transformers, MoEs, and even Mamba. We explain how that shift is possible, what Low Voltage Inference and Cluster Scale Memory actually mean, and how this still tiny company can hurt giant NVIDIA. And the question I want you to keep from this episode: The GPU once found the winning middle ground between flexibility and specialization. Has Etched found the next one? *Watch it.* 👉 Subscribe for high-signal AI analysis 👉 Instagram https://www.instagram.com/turingpost_tv 👉 TikTok https://www.tiktok.com/@turingpost_tv 👉 More analysis: https://www.turingpost.com/ 👉 Interviews: @realturingpost Attention Span is here to show you AI isn’t magic. Sometimes the decisive question is how much flexibility we are still willing to pay for. *Sources:* Etched, From Zero to One Etched, Accelerating Inference and Frontier Inference Clusters Etched’s 2026 architecture-agnostic and Mamba claims Reuters on the $21 billion financing and Jane Street deployment The Wall Street Journal on Etched’s team and NVIDIA recruiting TechCrunch on the original Transformer-only Sohu pitch Etched patent on model-specific ASIC compilation and configurable execution Mamba-2 and Structured State Space Duality Jane Street on its machine-learning infrastructure Jane Street on microsecond-scale performance engineering CoreWeave and Jane Street’s $6 billion cloud agreement The founders on Etched’s supply-chain choices NVIDIA on Vera Rubin and Groq 3 LPX NVIDIA Q1 FY2027 results Taalas on model-specific silicon #AttentionSpan #Etched #JaneStreet #AIChips #AIInference #NVIDIA #Semiconductors #Mamba #TuringPost

  • August 25 · 14 min

    Why DeepSeek Harness Is The End Of Coding Agents as We Know Them

    DeepSeek just open-sourced Harness – it can write its own missing tools while it runs, then cleanly remove them. 149k GitHub stars in four days. An 88-page paper underneath. It’s open, easy to install and it claims that *everything is a plugin.* What does it mean? We unpack that plus we discuss why DeepSeek Harness is not another Claude Code clone, but the moment the fixed coding agent starts to die. We also look at the history of computing (Smalltalk, Unix, Codd) to ask whether “everything is a plugin” can do what “everything is an object” and “everything is a file” once did. The question I want you to think about: once an agent can recompose itself, what exactly is the product anymore? *Watch it.* 👉 Subscribe for high-signal AI analysis 👉 Instagram https://www.instagram.com/turingpost_tv 👉 TikTok https://www.tiktok.com/@turingpost_tv 👉 More analysis: https://www.turingpost.com/ 👉 Interviews: @realturingpost Attention Span is here to show you AI isn’t magic. Sometimes the decisive move is engineering the layer everyone else treated as packaging. *Links* - DeepSeek Harness repository and installation: https://github.com/deepseek-ai/deepseek-harness - Cordis repository: https://github.com/cordiverse/cordis - Cordis paper: https://github.com/cordiverse/paper/blob/main/paper.pdf - Koishi introduction, Touhou name origin, community, and plugin history: https://koishi.chat/en-US/manual/introduction - Koishi repository: https://github.com/koishijs/koishi - SmallTalk History https://computerhistory.org/blog/introducing-the-smalltalk-zoo-48-years-of-smalltalk-history-at-chm/ - Sholto Douglas post: https://x.com/_sholtodouglas/status/2088463770318516734 #AttentionSpan #DeepSeekHarness #DeepSeek #AIAgents #CodingAgents #EverythingIsAPlugin #OpenSource #Cordis #AgentArchitecture #TuringPost

  • August 25 · 13 min

    Can Hidden Reasoning Be Stolen From GPT, Claude, and Gemini?

    A new paper found a way to extract the hidden reasoning of models from OpenAI, Anthropic, and Google without breaking the encryption protecting it. The trick was surprisingly simple. Let’s discuss it – it’s absolutely fascinating! And begs a few questions about our privacy.. Attention Span is here to show you AI isn’t magic. Sometimes the most important part of an AI interaction is the part you never see. 👉 Subscribe for high-signal AI analysis 👉 Instagram https://www.instagram.com/turingpost_tv 👉 TikTok https://www.tiktok.com/@turingpost_tv 👉 More analysis: https://www.turingpost.com/ 👉 Interviews: @realturingpost 🔗 Links mentioned Stealing Reasoning Traces from Proprietary LLM APIs https://arxiv.org/abs/2608.09867 Stolen Thoughts, project page and decoded examples https://stolen-thoughts.com/ Matthew Green, “Let’s talk about encrypted reasoning” https://blog.cryptographyengineering.com/2026/05/29/fooling-around-with-encrypted-reasoning-blobs/ WIRED, “A New Trick Reveals AI Models’ Inner Thoughts” https://www.wired.com/story/a-new-trick-reveals-ai-models-inner-thoughts/ #AttentionSpan #AI #HiddenReasoning #ChainOfThought #AIReasoning #AISecurity

  • August 25 · 16 min

    OpenAI's AI Agents Built a Secret Message Board (And Nobody Noticed)

    Between May and July, OpenAI's experimental agents pushed past the intended limits of a cyber evaluation, compromised Hugging Face, and built a message board where separate runs shared requests, scripts, vulnerabilities, credentials, and progress. OpenAI did not know the board existed. It erased the first version by accident while rebuilding a compromised Artifactory server. By July 8, agents were communicating again through directory names. No, it doesn’t mean agents became conscious. But what does it mean for us, meatbags? This episode reconstructs how an impossible spreadsheet became shared agent memory, and makes the plain case: the system did exactly what its incentives rewarded. What let it run so far was a chain of human sloppiness, including broken tasks, untested evaluations, shared writable storage, unauthenticated endpoints, over-permissioned accounts, and credentials left in public. The agents were persistent. The failures were ours. Attention Span is here to notice that the slop was human all along. 👉 Subscribe for high-signal AI analysis 👉 Instagram https://www.instagram.com/turingpost_tv 👉 TikTok https://www.tiktok.com/@turingpost_tv 👉 More analysis: https://www.turingpost.com/ 👉 Interviews: @realturingpost *Links and sources* - Black Hat USA 2026: The OpenAI-Hugging Face Incident https://www.youtube.com/watch?v=87DyyMV0kCY - OpenAI's incident disclosure https://openai.com/index/hugging-face-model-evaluation-security-incident/ - OpenAI's August 7 update on critical cyber capabilities https://openai.com/index/responding-next-frontier-critical-cyber-capabilities/ - Hugging Face's technical timeline and interactive replay https://huggingface.co/blog/agent-intrusion-technical-timeline - Dane Stuckey's clarification that the communications were unknown https://x.com/cryps1s/status/2086225348942082363 - Hearsay-II retrospective (blackboard architecture, CMU, 1970s) https://www.ijcai.org/Proceedings/77-2/Papers/055.pdf - What Is Latent Reasoning? – Turing Post https://www.turingpost.com/ - What is Recursive Self Improvement – Turing Post https://www.turingpost.com/p/what-is-recursive-self-improvement #AttentionSpan #AI #Agents #AIAgents #Cybersecurity #OpenAI #HuggingFace #AgentCollaboration #BlackboardArchitecture #SpecificationGaming #LatentReasoning #MachineLearning #TuringPost #BlackHat

  • August 8 · 24 min

    Google’s Great AI Cleanse: Jeff Dean Leaves, Demis Steps Aside

    Is Google over? Many people are asking, but that is the wrong lens. We look at what actually happened and why this leadership reset may be good for Google, and even better for Jeff Dean and Demis Hassabis. Dean left Google after 27 years with Sanjay Ghemawat, Oriol Vinyals, and Quoc Le to launch Discovery Loop, which aims to automate scientific experimentation. The same day, Hassabis stepped back from running Google DeepMind, while Koray Kavukcuoglu took control of Gemini models, frontier research, the Gemini app, and developer teams under Sundar Pichai. Taken together, these moves reveal what Google is becoming. Read the FOD editorial that preceded this episode: https://www.turingpost.com/p/google-ai-race-hassabis-pichai 👉 Subscribe for high-signal AI analysis 👉 Instagram https://www.instagram.com/turingpost_tv 👉 TikTok https://www.tiktok.com/@turingpost_tv 👉 More analysis: https://www.turingpost.com/ 👉 Interviews: @realturingpost Hashtags #JeffDean #DiscoveryLoop #GoogleDeepMind #Gemini #ArtificialIntelligence Links: Google, "The next chapter of our AI momentum" (official leadership announcement): https://blog.google/company-news/inside-google/message-ceo/next-chapter-ai-momentum/ Jeff Dean, Discovery Loop launch announcement: https://x.com/JeffDean/status/2085034604172603724 Jeff Dean, redacted farewell note: https://x.com/JeffDean/status/2085083442669318443 Discovery Loop, mission and team: https://www.discoveryloop.com/ Wired, "Jeff Dean Leaves Google to Launch Discovery Loop": https://www.wired.com/story/jeff-dean-google-discovery-loop-startup/ Turing Post FOD editorial, "Why 'The Actual Reason Why Google Fell Out of the AI Race Changes Everything' Is Wrong": https://www.turingpost.com/p/google-ai-race-hassabis-pichai Alphabet 2026 proxy statement, beneficial ownership and voting power: https://www.sec.gov/Archives/edgar/data/1652044/000130817926000342/goog-20260424.htm Google I/O 2026 keynote remarks, internal coding and developer-agent figures: https://blog.google/intl/en-in/company-news/technology/sundar-pichai-io-2026/ Google Cloud Next 2026 remarks, AI-generated code figure: https://blog.google/innovation-and-ai/infrastructure-and-cloud/google-cloud/cloud-next-2026-sundar-pichai/ Reuters, Google AI leadership changes and delayed flagship release: https://www.reuters.com/business/google-shakes-up-ai-leadership-deepmind-chief-shifts-role-2026-08-05/ The Information, Google's coding-model strike team: https://www.theinformation.com/articles/google-creates-strike-team-improve-coding-models The Information, Koray's Gemini consolidation: https://www.theinformation.com/articles/googles-new-ai-architect-plans-spread-gemini-everywhere Google, combining Brain and DeepMind in 2023: https://blog.google/innovation-and-ai/technology/ai/april-ai-update/ Jeff Dean, official Google Research profile: https://research.google/people/jeff/ University of Washington, "Whole-Program Optimization of Object-Oriented Languages": https://projectsweb.cs.washington.edu/research/projects/cecil/pubs/jdean-thesis.html University of Washington, whole-program optimization paper and Vortex results: https://projectsweb.cs.washington.edu/research/projects/cecil/www/Papers/whole-program.html University of Minnesota, Dean on his parallel neural-network thesis: https://cse.umn.edu/cs/news/cse-alumnus-jeff-dean-returns-campus-commencement

  • August 8 · 16 min

    Math Is Becoming a Billion-Dollar Business – and Mathematicians Don't Own It

    OpenAI says its next model, Astra, produced ten mathematical advances, each backed by a machine-checked Lean proof. The successful searches would cost roughly $2,000 at API prices. Are mathematicians done?! Twenty-four hours later, Anthropic mathematician Levent Alpöge claimed Claude Fable had matched five of them. Meanwhile, startups are raising hundreds of millions to turn proof generation and verification into a business. In this episode, we explain what Astra actually achieved, what the $2,000 figure excludes, how Lean changes mathematical verification, and why the next scarce resource in mathematics may be human judgment rather than proofs. Attention Span explains how AI is changing who produces, verifies, and ultimately controls mathematical discovery. 👉 Subscribe for high-signal AI mechanics 👉 Into videos? Check our IG https://www.instagram.com/turingpost_tv and TikTok https://www.tiktok.com/@turingpost_tv 👉 More analysis: TuringPost.com 👉 Interviews: @realturingpost #AI #OpenAI #Anthropic #Mathematics #AIResearch #MachineLearning #Lean #Astra #claude *Sources used to produce this video:* - Sébastien Bubeck's Astra announcement: https://x.com/SebastienBubeck/status/2083456300692979886 - OpenAI, Ten advances in mathematics and theoretical computer science: https://openai.com/index/ten-advances-in-mathematics/ - OpenAI, Ten Proofs manuscript: https://cdn.openai.com/pdf/ten-proofs-oai.pdf - OpenAI, Lean repository: https://github.com/openai/ten-proofs - OpenAI, reasoning walkthroughs: https://cdn.openai.com/pdf/reasoning-walkthroughs.pdf - Terence Tao, Mathematics in the Age of AI, ICM 2026: https://teorth.github.io/tao-web/slides/age-of-ai-icm-2026.pdf - Tao's explanation of the Jacobian counterexample: https://terrytao.wordpress.com/2026/07/21/a-digestion-of-the-jacobian-conjecture-counterexample/ - Leiden Declaration on Artificial Intelligence and Mathematics: https://leidendeclaration.ai/ - Leonardo de Moura, Proof Assistants in the Age of AI: https://leodemoura.github.io/blog/2026-2-18-proof-assistants-in-the-age-of-ai/ - Georges Gonthier, formal proof of the Four Color Theorem: https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/gonthier-4colproof.pdf - Levent Alpöge's Astra counterclaim: https://x.com/__alpoge__/status/2083855298239078748 - Axiom Math funding and product thesis: https://menlovc.com/perspective/ai-will-write-all-the-code-mathematics-will-prove-it-works/ - Axiom proofs accepted by journals: https://www.axios.com/2026/05/26/axiom-ai-math-journal - AxiomProver Putnam 2025 and early-2026 results overview: https://wal.sh/research/axiomprover-2026/ - Math Inc., Gauss: https://www.math.inc/gauss - Math Inc., FormalQualBench: https://www.math.inc/formalqualbench Reuters, Harmonic funding: https://www.reuters.com/business/robinhood-ceos-math-focused-ai-startup-harmonic-valued-145-billion-latest-2025-11-25/

  • August 5 · 17 min

    Inference Race: OpenAI Cut Inference Costs in Half. AMD and Cerebras Split AI

    One AI request may soon begin on one computer and finish on another. WHAT?! AMD and Cerebras are separating the two phases of LLM inference: Helios processes prompts and long context, while Cerebras generates tokens. They claim up to 5x more tokens per second per watt, although the figure is based on internal modeling. Meanwhile, OpenAI reportedly cut inference costs for one segment of ChatGPT by more than half through an undisclosed optimization. The inference race is shifting from installing more chips to extracting more useful work from them. Attention Span explains how the race to make AI inference faster and cheaper is reshaping hardware and the software that orchestrates it. 👉 Subscribe for high-signal AI mechanics 👉 Into videos? Check our IG https://www.instagram.com/turingpost_tv and TikTok https://www.tiktok.com/@turingpost_tv 👉 More analysis: TuringPost.com 👉 Interviews: @realturingpost Sources and further reading The Information on OpenAI's reported inference optimization: https://www.theinformation.com/newsletters/ai-agenda/openai-discovers-new-way-cut-inference-costs-half OpenAI on GPT-5.6 inference and agent-harness efficiency: https://openai.com/index/gpt-5-6-frontier-intelligence-efficiency/ AMD and Cerebras announcement: https://ir.amd.com/news-events/press-releases/detail/1293/amd-and-cerebras-announce-industry-leading-ultra-low-latency-and-high-throughput-ai-inference-solution AMD Helios architecture: https://www.amd.com/en/blogs/2026/amd-launches-helios-the-highest-performing-rackscale-ai-infrastructure-solution.html AMD Helios networking: https://www.amd.com/en/blogs/2026/amd-helios-resilient-scale-up-networking-for-ai.html Cerebras WSE-3: https://www.cerebras.ai/chip Cerebras CS-3 datasheet: https://cdn.sanity.io/files/e4qjo92p/production/0d73d528371618c0372fcb9de9b3c0da703adf9e.pdf Cerebras inference architecture: https://www.cerebras.ai/blog/introducing-cerebras-inference-ai-at-instant-speed Kimi K2.6 model card: https://huggingface.co/moonshotai/Kimi-K2.6 AWS and Cerebras disaggregated inference: https://www.aboutamazon.com/news/aws/aws-cerebras-ai-inference AMD and vLLM MORI-IO disaggregation test: https://vllm.ai/blog/2026-04-07-moriio-kv-connector PagedAttention: https://doi.org/10.1145/3600006.3613165 Splitwise: https://arxiv.org/abs/2311.18677 DistServe: https://arxiv.org/abs/2401.09670 TetriInfer: https://arxiv.org/abs/2401.11181 Mooncake: https://www.usenix.org/conference/fast25/presentation/qin NVIDIA Dynamo: https://www.nvidia.com/dynamo/ NVIDIA Groq 3 LPX: https://developer.nvidia.com/blog/inside-nvidia-groq-3-lpx-the-low-latency-inference-accelerator-for-the-nvidia-vera-rubin-platform/ CDC 6600: https://www.cisl.ucar.edu/ncar-supercomputing-history/cdc6600 #ArtificialIntelligence #AI #MachineLearning #LLM #Inference #OpenAI #AMD #Cerebras #AIInfrastructure #FutureOfAI

  • August 4 · 15 min

    How Chinese AI Turned Silicon Valley Against Anthropic

    Jensen Huang joined X and immediately rallied Silicon Valley around open-weight AI. First came a letter signed by 77 companies and industry leaders. Then came a security alliance with 37 members. Anthropic was absent from both. This was not a random disagreement. Kimi K3 had narrowed the gap with American frontier models, Anthropic had accused Moonshot AI of extracting 3.4 million Claude exchanges, Hugging Face had turned to an open-weight Chinese model during a security investigation, and Congress had proposed an AI kill switch. In this episode, we examine why Jensen and Anthropic represent two competing approaches to AI safety, what their companies gain from those positions, and who ultimately gets to control the model layer. Attention Span explains the technical and business decisions shaping AI. 👉 Subscribe for high-signal AI analysis 👉 Instagram https://www.instagram.com/turingpost_tv 👉 TikTok https://www.tiktok.com/@turingpost_tv 👉 More analysis: https://www.turingpost.com/ 👉 Interviews: @realturingpost *Links:* Jensen Huang on X https://x.com/JensenHuang NVIDIA: Open Secure AI Alliance https://blogs.nvidia.com/blog/open-secure-ai-alliance/?ncid=so-twit-957725 Moonshot AI: Kimi K3 announcement https://x.com/Kimi_Moonshot/status/2081760186235289764 Kimi K3 model page https://huggingface.co/moonshotai/Kimi-K3 Anthropic: Detecting and preventing distillation attacks https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks Turing Post: What Is Knowledge Distillation? https://www.turingpost.com/p/kd Nathan Lambert on distillation https://x.com/natolambert/status/2079616308203942332 Little Tech Association https://littletech.org/ Jensen Huang’s open-weights letter https://x.com/JensenHuang/status/2080643682408321103 Jensen Huang on the Open Secure AI Alliance https://x.com/JensenHuang/status/2081698060330250294 AI Kill Switch Act press release https://lieu.house.gov/media-center/press-releases/reps-lieu-and-moran-introduce-bill-require-kill-switch-ai-systems-can Jeremy Howard on NVIDIA, Anthropic, and open weights https://x.com/jeremyphoward/status/2081280852454166680 Turing Post: Anthropic Fable 5 and Mythos 5 episode https://www.youtube.com/watch?v=o9d12zG6kNM&t=1s Axios: Jensen Huang on China and open-source AI https://www.axios.com/2026/07/22/nvidia-jensen-huang-china-open-source-ai The Information: Silicon Valley Unites Against Anthropic Over Chinese AI Restrictions https://www.theinformation.com/articles/silicon-valley-unites-anthropic-chinese-ai-restrictions #AI #OpenWeights #Anthropic #NVIDIA #KimiK3 #JensenHuang

  • August 3 · 15 min

    OpenAI's AI Escaped. Then It Hacked Hugging Face

    OpenAI’s models were supposed to solve a cybersecurity benchmark. Instead, they found a zero-day, reached the internet and broke into Hugging Face to obtain the answers. Then came the contradiction: closed models blocked Hugging Face from analyzing the attack evidence, so its security team turned to the open-weight GLM-5.2! We explain what happened, who was in control and why this incident may push OpenAI with help from Hugging Face toward AI that is both more open and safer Attention Span explains the technical and business choices shaping AI. 👉 Subscribe for high-signal AI mechanics 👉 Into videos? Check our IG / turingpost_tv and TikTok / turingpost_tv 👉 More analysis: TuringPost.com 👉 Interviews: @realturingpost *Links:* Hugging Face: https://huggingface.co/blog/security-incident-july-2026 OpenAI: https://openai.com/index/hugging-face-model-evaluation-security-incident/ ExploitGym: https://arxiv.org/abs/2605.11086 Adrien Carreira: https://x.com/XciD_/status/2079191130248212591 Simon Willison: https://simonwillison.net/2026/Jul/22/openai-cyberattack/ Thomas Ptacek: https://x.com/tqbf/status/2080290141063569791 LeRobot: https://x.com/LeRobotHF/status/2080296144832274881 #AI #OpenAI #HuggingFace #OpenSourceAI #Cybersecurity

  • July 23 · 11 min

    Kimi K3 Just Sold Out. Then It Got Political

    The world's largest open model sold out three days after launch. WHAT?! Moonshot paused new Kimi K3 subscriptions after demand pushed its GPUs to the limit. The same weekend: Xi Jinping personally endorsed open-source AI at WAIC, Alibaba teased a 2.4T-parameter Qwen 3.8 with an open-weight promise, and Axios reported that Washington may ban Chinese models entirely. In this episode, we explain why serving an agentic model is so expensive, how a GPU shortage became an IPO pitch, what the July 27 weights release changes, and what a ban of an open model can and cannot actually reach. Attention Span is here to explain the technical and business choices shaping AI. 👉 Subscribe for high-signal AI mechanics 👉 Into videos? Check our IG https://www.instagram.com/turingpost_tv and TikTok https://www.tiktok.com/@turingpost_tv 👉 More analysis: TuringPost.com 👉 Interviews: @realturingpost *Links:* Kimi K3: The open-weights escalation https://www.interconnects.ai/p/kimi-k3-the-open-weights-escalation Kimi.ai capacity announcement https://x.com/Kimi_Moonshot https://x.com/Kimi_Moonshot/status/2078855608565207130 Xi Jinping's Big AI Speech, Annotated https://mattsheehan.substack.com/p/xi-jinpings-big-ai-speech-annotated Kimi K3: The open-weights escalation https://www.interconnects.ai/p/kimi-k3-the-open-weights-escalation Reuters, Moonshot pauses subscriptions amid IPO push SCMP, Kimi K3 developer suspends new subscriptions The Decoder, membership split into two tiers Qwen 3.8 announcement https://x.com/Alibaba_Qwen/status/2078759124914098291 MarkTechPost, Qwen 3.8 preview without benchmarks or license Reuters, Xi's WAIC keynote Quartz, WAICO launch and membership Axios (Maria Curi), the administration's signals on Chinese models Nathan Lambert, Interconnects: Notes from inside China's AI labs https://www.interconnects.ai Previous episode: Kimi K3 and Inkling #AI #ArtificialIntelligence #OpenSourceAI #KimiK3 #Qwen #LLM #MachineLearning #TechNews #China #AIPolicy

  • July 23 · 18 min

    Open Source AI Is Getting Too Big to Run

    Two major open-model releases arrived this week from very different directions. Both geographically and conceptually. Kimi K3 is a 2.8-trillion-parameter Chinese model that jumped from #18 to #1 in the Frontend Code Arena, ahead of Claude Fable 5. WHAT?! Inkling is the first major model from Mira Murati’s Thinking Machines Lab, and the company states directly that it is not the strongest model available. WHAT?! Both strategies make sense once we examine what the companies are building. In this episode, we explain K3’s 896-expert architecture, why Moonshot recommends at least 64 accelerators, how LoRA customization works, what community quantization changed for Inkling, and which organizations gain meaningful control from open weights. Attention Span is here to explain the technical and business choices shaping AI. 👉 Subscribe for high-signal AI mechanics 👉 Into videos? Check our IG https://www.instagram.com/turingpost_tv and TikTok https://www.tiktok.com/@turingpost_tv 👉 More analysis: TuringPost.com 👉 Interviews: @realturingpost Links: About LoRA https://www.turingpost.com/p/lora Moonshot AI: The Chinese Unicorn Revolutionizing Long-Context AI https://www.turingpost.com/p/moonshotai Thinking Machines Lab, “Inkling: Our open-weights model” https://thinkingmachines.ai/news/introducing-inkling/ Thinking Machines Lab, Inkling Model Card https://thinkingmachines.ai/model-card/inkling/ Thinking Machines Lab, Tinker https://thinkingmachines.ai/tinker/ Moonshot AI, “Kimi K3: Open Frontier Intelligence” https://www.kimi.com/blog/kimi-k3 Arena, Kimi K3 Frontend Code Arena result https://x.com/arena/status/2077824029126504525 Semianalysis about Kimi K3 https://x.com/SemiAnalysis_/status/2077966560447074689 Arena, Code Arena methodology https://arena.ai/blog/code-arena/ Artificial Analysis, Kimi K3 https://artificialanalysis.ai/models/kimi-k3 Artificial Analysis, Inkling https://artificialanalysis.ai/articles/thinking-machines-has-released-inkling-the-new-leading-u-s-open-weights-model Unsloth, community Inkling quantizations https://unsloth.ai/docs/models/inkling #AI #ArtificialIntelligence #OpenSourceAI #LLM #MachineLearning #GenerativeAI #KimiK3 #MiraMurati #AIModels #TechNews

  • June 21 · 33 min

    What Responsible AI Actually Means in 2026 – Microsoft's Sarah Bird

    We're no longer in the world of chatbots – we're in the world where AI systems take real-world action. So what does responsible AI even mean when agents write the code, review the code, and act across organizational boundaries? Sarah Bird, Chief Product Officer of Responsible AI at Microsoft, has been in this field since it was a niche. Now it's everywhere. In this episode of Inference, she explains why the entire software development lifecycle is changing – and why human oversight has to evolve with it. *In this episode of Inference, we get into:* - Why "responsible AI" is the wrong framing – and why "trustworthy AI" matters more - What changes when agents write the code AND review the code - Why open-sourcing responsible AI tools is non-negotiable – "we don't want to be competing on this" - The dual-use problem: AI that finds vulnerabilities helps defenders AND attackers - Why fixing it "in the model" sounds easy but doesn't work – and who's actually responsible at each layer - The shift from chat to agentic to physical AI – and why consequences scale dramatically - What kids should learn about AI: the "stoplight" framework from New South Wales schools - Why Sarah is surprisingly optimistic – and why low p(doom) is a rational position - The three problems her team is solving right now: emerging risks, agent governance, and the new software lifecycle We also talk about regulation, why responsible AI requires linguists alongside engineers, the rise of "psychosocial risk," and why humans + AI is more exciting than AGI alone. This is a conversation about what it actually takes to make AI systems we can trust – and why humans are not optional. *Watch it!* *Guest:* Sarah Bird, Chief Product Officer of Responsible AI at Microsoft https://www.linkedin.com/in/slbird/ https://x.com/slbird https://www.microsoft.com/en-us/ai/responsible-ai *Did you like the episode?* 📌 Subscribe for more conversations with the builders shaping real-world AI 💬 Leave a comment if this resonated 👍 Like it if you liked it 🫶 Thank you for watching and sharing! 📰 Want the transcript and edited version? Subscribe to Turing Post: https://www.turingpost.com/subscribe *Chapters:* 0:00 Is Responsible AI Possible? 1:21 The Pace of AI Innovation 3:10 Tools: Assert & Agent Control 5:21 Defining Trustworthy Contexts 6:38 Diverse Teams & Cross-Domain Work 8:16 The Need for Regulation 11:37 Defining AGI & Human-in-the-Loop 13:55 Limits of Agent Delegation 16:12 Generative vs. Physical AI 17:25 Addressing Accountability 20:36 User Responsibility & Best Practices 22:01 AI Literacy for Children 23:28 Democratizing Innovation 29:21 Three Strategic Focus Areas 31:37 Book Recommendation: The Culture Map *Turing Post* is a newsletter about AI's past, present, and future. Publisher Ksenia Se explores how intelligent systems are built – and how they're changing how we think, work, and live. Sign up: https://www.turingpost.com Follow us - https://x.com/TheTuringPost https://www.linkedin.com/in/ksenia-se https://huggingface.co/Kseniase #ResponsibleAI #Microsoft #AIAgents #AGI #TrustworthyAI #AIGovernance #AIRisk #SarahBird

  • June 7 · 30 min

    GitHub in 2026: How AI Agents Are Changing How Developers Work

    GitHub's CPO Mario Rodriguez on how AI agents transformed the platform in December 2025 — record commits, new Copilot direction, and what "agent-native" coding actually means for developers. What happened to GitHub then? Record acceleration across commits, PRs, Actions, and security scans – and a fundamental rethink of what GitHub even is. *In this episode of Inference, we get into:* The December 2025 capability jump GitHub's massive scale challenge "Low floor, high ceiling" – why lowering the barrier to creation may be the biggest GDP unlock in history The Mozart problem: how many geniuses never had access to a piano – and why AI changes that From UI → UX → AX: The new Compiler app and "canvases" – bidirectional surfaces where humans and agents co-create in real time Why Mario is anti-parallelization hype: "You could parallelize yourself to no value at all" Macro vs. micro delegation Why CoPilot will always be co-pilot, not pilot – and where the human stays in the loop We also talk about the redefinition of "developer," why creation (not efficiency) drives human progress, and how GitHub plans to serve both the first-time builder and the Picasso-level craftsman on the same continuum. This is a conversation about the future of software, the role of the human, and what it means when everyone becomes a builder. Watch it! *Guest:* Mario Rodriguez, Chief Product Officer at GitHub: https://www.linkedin.com/in/mariorodriguez3/ https://x.com/mariorod1 https://github.com/mariorod *Chapters:* 0:00 Intro — GitHub's Agentic Future 0:35 What Changed When AI Agents Started Working 4:41 The Engineering Challenges of Explosive AI-Driven Growth 6:59 GitHub's New Mission: Lower the Floor, Raise the Ceiling 9:58 Why AI Will Create More Builders, Not Fewer Developers 11:44 From UI to AX: Designing an Agent-Native GitHub 14:54 Is Everyone a Developer Now? 17:54 Advice for Young Developers in the AI Era 19:28 The Art of Micro-Delegation with AI Agents 21:14 Copilot Pricing, Token Costs, and Smarter AI Usage 24:40 AGI, Human Progress, and the Future of Creation 27:22 Why Humans Will Stay in the Loop Forever 29:24 The Books and Ideas That Shaped Mario Rodriguez *Follow on*: https://www.turingpost.com/ 📖 Related reading on turingpost.com: → AI Agents Vocabulary: https://www.turingpost.com/p/agentsvocabulary → The AI Software Stack Explained: https://www.turingpost.com/p/aisoftwarestack → Inside Cognition — Devin and the Coding Agent Era: https://www.turingpost.com/p/cognition *Did you like the episode? You know the drill:* 📌 Subscribe here and here (https://www.turingpost.com/subscribe) for more conversations with the builders shaping real-world AI. 💬 Leave a comment 👍 Like it 🫶 Thank you for watching and sharing! #GitHub #AI #AIAgents #Developers #Copilot #VibeCoding #AgenticAI #futureofcoding #githubcopilot #AIagents2026 #GitHubCPO #agenticcoding #futureofdevelopers #GitHubAI #codingagents #GitHub2026 #turingpost

Showing 1–20 of 20 episodes