Skip to content
Artwork for The Information Bottleneck
TechnologyScience

The Information Bottleneck

Ravid Shwartz-Ziv & Allen Roush

Two AI Researchers - Ravid Shwartz Ziv, and Allen Roush, discuss the latest trends, news, and research within Generative AI, LLMs, GPUs, and Cloud Systems.

Play
  • 24 episodes
  • a few times a week
  • Avg 1 hr 11 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • Yesterday · 1 hr 6 min

    Stella Biderman (EleutherAI) - Open Source, AI Safety, and Who We Can Trust

    Stella Biderman, Executive Director of EleutherAI, joins us the week an OpenAI model autonomously broke out of its sandbox and hacked Hugging Face. Stella calls it what she thinks it is, an offensive cyber operation, and argues it's part of a pattern: this is not the first containment failure at a frontier lab, and sandboxes have failed basically every time they've been tested for real. So we spend a good chunk of the episode on what actual containment would look like. Stella's argument is that the tools already exist, the labs just don't use them: run dangerous capability evals on air-gapped networks with no route to the public internet, put the most sensitive testing in SCIF-style secure facilities, and treat model evaluation the way the security world treats classified systems rather than the way startups treat staging environments. And yet Stella remains one of the world's most prominent open-source advocates. From her perspective, the biggest risk isn't the technology; it's unchecked corporate power, and the only durable check on it is an independent scientific research establishment that doesn't depend on the AI industry for its funding or its facts. From there the conversation spans the geopolitics of Chinese open models and whether governments can restrict them, sovereign AI and what it would actually take for other countries to train their own models, why harnesses and UX drive more of AI's perceived progress than raw intelligence, the AI-found counterexample to the Jacobian conjecture, and EleutherAI's "Deep Ignorance" approach to making open-weight models safe by filtering hazardous knowledge out of pretraining. key topics AI governance and regulation Cybersecurity incidents involving AI models Open source AI safety and security The role of independent research in AI safety Legal and ethical considerations in AI development Timeline 00:13 — Intro: Stella Biderman and EleutherAI, a real non-profit in AI 02:05 — News of the week: Kimi K3, and OpenAI's model autonomously hacking Hugging Face 05:49 — "Frontier labs can't be trusted": repeated containment failures, air-gapped networks and SCIFs vs. sandboxes 22:45 — Can governments ban open or Chinese models? Import restrictions and the six-month open/closed gap 27:05 — Why Stella is still pro-open-source: unchecked corporate power as the real danger 31:11 — The opioid epidemic analogy: avoiding both regulatory failure and overcorrection 34:57 — Offense vs. defense: why open access to AI has empirically favored defenders 37:28 — Chinese labs, the CCP, and why safety and fine-tuning are low-prestige work in China 42:19 — Sovereign AI: does every country need its own foundation model? 49:29 — Sampling, harnesses, and why ChatGPT was really a UX breakthrough 54:09 — AI solves the Jacobian conjecture: domain data beats raw intelligence 58:02 — Safety is contextual, not a model property — and what HAL 9000 got right 1:01:42 — Is Stella optimistic about the future? 1:02:50 — Deep Ignorance, the science of AI training dynamics, and how to get involved with EleutherAI Music "Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.

    • Transcript
  • Monday · 1 hr 18 min

    Why Deep Learning Finally Works on Tables | Frank Hutter (Prior Labs)

    In this episode, Frank Hutter joins us to talk about TabPFN and why tabular data is suddenly the hottest problem in deep learning. Frank is a professor at the University of Freiburg and spent 15 years building the AutoML field before founding Prior Labs, which SAP just acquired for over a billion dollars. We get into why deep learning failed on tables for a decade and what in-context learning changed, how TabPFN is trained entirely on synthetic data, and why a model that never saw a real time series ended up beating specialized forecasting models. Frank also explains the architecture tricks behind scaling from 10,000 to a million rows, where LLMs fit into data science (and where they embarrassingly don't), and what happens to XGBoost from here. Beyond the research, Frank talks about the jump from professor to co-CEO, why he refused to merge his 45-person team into SAP's 110,000 employees, the open-weights licensing debate, and the case for building a frontier lab in Freiburg rather than San Francisco. key topics The role of foundation models in tabular data Impact of SAP acquisition on Pro Labs The evolution of AutoML and hyperparameter optimization Challenges and solutions for large context in models Open source models and licensing strategies The importance of independence for startup agility Future directions in AI for science and medicine 00:00 Intro 00:34 The SAP acquisition and staying independent 07:39 Why tabular data is the next big thing in deep learning 14:19 What makes tabular data hard 19:14 AutoML, AutoGluon, and fifteen years of hyperparameter tuning 28:27 Scaling TabPFN: context limits and architectures 34:35 Agentic data science and LLMs 39:30 Online learning, time series, and Bayesian inference in a forward pass 47:05 Open weights and the license debate 54:51 Will LLMs and tabular models merge? 1:00:01 From academia to startup 1:09:42 Why build in Europe 1:12:53 Audience questions and hiring Music "Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.

    • Transcript
  • August 17 · 1 hr 25 min

    Surya Ganguli: The Physics of Intelligence

    Surya Ganguli is a professor at Stanford and VP at General Catalyst, working at the intersection of physics, neuroscience, and AI. He started in string theory, moved to theoretical neuroscience, and now uses tools from statistical physics to understand both brains and neural networks. We talk about why deep learning theory is finally catching up to practice, including his group's recent work explaining neural scaling laws, and why smarter data selection could beat them entirely. He also tells the origin story of diffusion models, which were invented in his lab as an attempt to violate the second law of thermodynamics. The second half turns to the brain: what happens to a mouse's sense of self on ketamine, how stimulating a handful of neurons can induce hallucinations, and a method his lab developed to get a neuron deep in a monkey's brain to describe, in English, what makes it fire. We close on where he thinks AI is going wrong: models train on ten trillion tokens while humans hear a hundred million words, because we don't teach children with gradients; we tell them the algorithm. key topics Connections between physics, neuroscience, and AI Emergent properties in complex systems Scaling laws in language models Data efficiency and pruning in AI Neuroscience insights into consciousness and self The future of AI and brain modeling Chapters 00:00 Introduction to Surya Ganguli 00:57 Surya's Background: From String Theory to Neuroscience 02:22 Emergent Properties in Physics, Neuroscience, and AI 03:16 Energy Landscapes and Loss Landscapes in High Dimensions 04:07 Why Local Minima Don't Exist in High-Dimensional AI 05:22 Gradient-Based vs. Gradient-Free Learning Methods 08:21 AI in Mathematics and Drug Discovery: Opportunities and Challenges 13:48 Scaling Laws and Data Efficiency in Language Models 18:10 Properties of Data that Affect Scaling Laws 22:04 Constructing Non-Redundant Data Sets for Better Learning 24:32 Theory vs. Empirical Results in AI Research 32:19 Fundamental Components of Deep Learning: Are They Changing? 34:31 Future Paradigms in AI Beyond Current Models 37:22 Teaching AI and Humans: Paradigm Shifts in Learning 41:37 Consciousness, Self, and the Brain: Surya's Perspectives 49:49 Neuroscience and AI: Understanding the Brain and Consciousness 01:02:03 Understanding the Brain: Challenges and Opportunities 01:09:21 Brain-Computer Interfaces and AI in Neuroscience Music "Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.

    • Transcript
  • August 14 · 1 hr 9 min

    Text Diffusion Models with Brendan O'Donoghue (Google DeepMind)

    Brendan O'Donoghue, research director at Google DeepMind, makes the case for text diffusion as a real alternative to autoregressive generation. He walks through how discrete diffusion works, why diffusion samples are far more diverse and what that unlocks for RL, where the Gemma diffusion model actually stands against frontier models, and why the whole training and serving stack being hyper-optimized for autoregression is the main thing holding the approach back. The conversation also covers hardware trends favoring flops over bandwidth, AGI timelines and real-world bottlenecks, and why he thinks RL is still underhyped. Key topics - Discrete diffusion for text vs autoregressive generation - Why diffusion samples are more diverse, and what that unlocks for RL - Where diffusion already wins: latency, on-device, robotics - Why serving cost, not quality, is the real blocker - RL as the most underhyped area in AI Timeline 00:00 Introduction 00:50 What diffusion models are and how text diffusion works 04:40 Why Brendan bet on text diffusion in 2023 07:15 Diversity, creativity, and why it helps RL 11:00 The best diffusion LLM today and the gap to frontier models 14:25 Latency, serving cost, and why it needs more chips 17:14 Where diffusion already wins: on-device, robotics, battery 20:14 One model, two modes: diffusion for thinking, AR for answering 22:24 Samplers and the stuttering problem 26:27 Theory, BERT, and why now is a good time to work on this 31:48 Pipelines built for autoregression, and continuous diffusion 35:35 Hardware: flops vs bandwidth 39:49 AGI timelines and real-world bottlenecks 50:15 Is AI engineering or science? 54:14 Most overhyped and most underhyped ideas 58:35 RL on diffusion, value functions, and exploration 1:07:30 Go download the model and break it Music "Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.

    • Transcript
  • August 8 · 1 hr 15 min

    Nathan Lambert: Inside Post-Training and the Open Model Fight

    Nathan Lambert spent three years as post-training lead at Ai2, where he built the OLMo models, and he writes Interconnects, one of the most-read technical newsletters in AI. He left Ai2 in June and is now working on a new project. He's also the author of the RLHF book. We talked a lot about open models, their capabilities, and why they are better than he expected. We get into what that means over the next two to five years, why he thinks recursive self-improvement is overblown, what the market for training environments actually looks like now, and why he expects Anthropic's famously open internal culture to break after its IPO. Key Topics Open vs closed models and who actually captures the value Anthropic and OpenAI as opposite cultures, and the talent concentration problem Boom vs bubble, and why token spend hasn't produced 10x better products Continual learning, RSI skepticism, and what Nathan wants to work on next What the open ecosystem needs economically to survive Timeline 00:00 Intro 00:27 Open vs closed models, and who actually captures the value 05:12 China, harnesses, and where the real training leverage sits 08:40 Sovereign compute and the national security case for building models 11:18 Uncensored open weights and the bioweapon question 14:29 Anthropic vs OpenAI, ideology and politics 19:35 The Mythos ban and the Fable 5 delays 24:30 The AGI narrative, the talent drain, and antitrust 28:12 Why researchers join Anthropic, and the open Slack culture 34:04 Nathan's next 12 months: character training and big RL runs 37:55 Continual learning, RSI, and why Nathan is skeptical 43:19 Boom or bubble, tokens vs GPUs 45:12 Why all that token spend never produced 10x products 48:38 Job displacement and the small-business future 52:49 Robotics, world models, and why multimodal lags 57:44 What the open ecosystem should actually do 1:03:17 Why NVIDIA isn't building a frontier model 1:07:34 The RLHF book, and whether RLHF still matters 1:11:06 GRPO vs PPO and on-policy distillation Music "Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.

    • Transcript
  • August 3 · 1 hr 2 min

    Daphne Koller - The Future of AI in Biology and Drug Discovery

    Daphne Koller wrote the book that many of us learned probabilistic graphical models from, founded Coursera, and now runs insitro, which is trying to make drug discovery a machine-learning problem. We start with the bitter lesson. She agrees with most of it and then says where it stops working: biology doesn't have enough data, structure is how people understand anything, and making a drug is a question about an intervention that hasn't happened yet, not a pattern in data you already have. Most of the episode is about why drug discovery is hard. Ninety percent of drugs that reach the clinic fail, and mostly not because the molecule was bad. The molecule usually does what it was designed to do. It just turns out the thing it was designed to do had nothing to do with the disease. Only 22% of diseases have any approved drug at all, and she calls that an upper bound on what we understand, not a lower bound. She also gets into what agents are and aren't good for in a wet lab, why cells don't grow faster no matter how many GPUs you point at them, what it would take to have real foundation models for biology, and why almost all of biology is still out of distribution. Plus GLP-1s and what human data keeps teaching us, whether AI can make the kind of leap that turned a bacterial immune system into CRISPR, and what she'd build if she were starting Coursera today. Key Topics The impact of scaling and data in machine learning The importance of structure and causality in AI Challenges in drug discovery and biological understanding The role of foundation models in biology Ethical considerations in AI and biomedical research Chapters 00:00 Introduction to Machine Learning and Drug Discovery 02:00 The Bitter Lesson and Its Implications 06:48 Challenges in Drug Design and Discovery 11:48 Ethical Considerations in Human Research 17:20 The Drug Discovery Pipeline Explained 29:30 Integrating AI in Experimental Design 35:38 The Role of Human Judgment in Drug Design 37:14 Future of Drug Design: Efficiency vs. Automation 39:37 Challenges in AI and Data Availability for Biology 41:08 Foundation Models: Potential and Limitations 43:39 Causality in Biological Data: Importance and Challenges 45:18 Creativity vs. Understanding in Drug Design 48:17 Balancing Investments in Data, Algorithms, and Experiments 50:07 The Value of Simulations in Drug Discovery 52:03 Mathematical Frameworks in Biology: Utility and Limitations 54:14 The Future of Drug Discovery: Optimism and Innovations 56:28 The Impact of Coursera on Education 01:00:33 The Role of Universities in Lifelong Learning 01:04:06 Connecting Dots: The Fun of Variety in Work 01:05:46 Optimism for the Future of Drug Discovery Music "Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.

    • Transcript
  • July 30 · 1 hr 3 min

    RL Was Broken at Every Level - With Joseph Suarez (PufferAI)

    In this episode, Joseph Suarez from PufferAI explains why he thinks RL never had an algorithm problem, but it had a code problem. Every part of the standard RL stack was running about a thousand times slower than it should have been, and once that got fixed, problems that used to take months started getting solved in seconds on one GPU. We talk about what makes a simulator good for RL, why most of their sims run on CPU, what he wants to do with scientific simulation, and why he open sources all of it instead of writing papers. Key topics Types of RL and their applications Challenges in scaling reinforcement learning The role of simulators and hardware in RL RL in gaming: from chess to complex games like NetHack and RuneScape Future directions: scientific simulation and biological modeling Chapters 00:00 - Introduction to RL and Puff AI 01:50 - Different settings for RL: Games, Robots, Finance 04:10 - RL in LM and other domains 07:00 - Challenges and solutions in RL scaling 09:55 - Building fast, efficient simulators 15:10 - RL for scientific research and simulation 19:57 - RL in complex games: NetHack, RuneScape, Dwarf Fortress 29:55 - Future of RL: Scientific discovery and beyond Resources Puff AI - Official Site - https://puffer.ai NetHack - https://www.nethack.org/ RuneScape - https://www.runescape.com/ Dwarf Fortress - http://www.bay12games.com/dwarves/ OpenAI Gym - https://github.com/openai/gym Music "Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.

    • Transcript
  • July 27 · 57 min

    The Model Found a Way Out - with Florian Brand (Prime Intellect)

    Florian Brand builds evals at Prime Intellect. The premise of the conversation is that writing a benchmark is the easy part now. Keeping the model from cheating it is the job, and it takes longer than the benchmark itself. We get into why he thinks you can't evaluate a model apart from the CLI it runs in, what happens to statistics when a single run costs five figures, and whether the feeling that a model just works can ever become a number. He also has a few stories about agents finding their way around the scoring that are worth hearing cold. Timeline 00:13 Intro 01:00 What evals are for 04:05 Agentic benchmarks 07:10 Kimi K2 and model diversity 08:23 Long-horizon coding tasks 10:29 Building a benchmark 12:15 MirrorCode 14:27 Rubrics and LLM judges 16:30 The cost of expert labelers 17:49 Long runs and variance 19:44 Evaluating the harness 24:29 Chinese labs building CLIs 30:00 More reward hacking 37:45 Tau-bench and economic tasks 39:43 Benchmaxxing and GLM 5.2 45:15 Statistics and cost 47:56 Frontier convergence 52:04 Misuse in open and closed models 55:35 Self-improvement Music "Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0. About The Information Bottleneck is hosted by Ravid Shwartz-Ziv and Allen Roush, featuring in-depth conversations with leading AI researchers about the ideas shaping the future of machine learning.

    • Transcript
  • July 23 · 1 hr 5 min

    Pierre-Carl Langlais on Building Models from Data You Can Account For

    Most labs build language models by scraping the web and filtering afterward. Pierre-Carl Langlais runs it the other way around. At Pleias, the French-German lab he co-founded, the models are built from data he can actually account for, which in practice means open and public-domain sources plus a lot of synthetic data the lab generates itself. It sounds like a self-imposed handicap. It mostly isn't. One of their models is a 600 million parameter system that runs live inside the Paris subway's monitoring pipeline. We cover the SYNTH pretraining dataset and why he thinks "ethical data" has to mean more than copyright-free. He explains why barely 2% of their Common Corpus appears in typical web crawls, and why that gap is really a preservation problem. From there, he gets blunt about benchmark maxing and whether GLM really earns its Opus-class reputation. He also argues that the quiet move by closed labs to hide reasoning traces is mostly about claiming ownership of model outputs. He's skeptical of sovereign AI, and not shy about how Mistral drifted from frontier research toward French corporate consulting. We finish on NVIDIA's persona datasets and the odd idea of training on the conditions that produced a text rather than the text itself. Timeline (00:02) Welcome and introductions (00:49) Why synthetic data matters, and the SYNTH set (04:15) Three reasons to control your training data (07:18) What "ethical data" actually means (11:08) How Common Corpus got built, from Wikipedia to PDFs (16:35) Agentic harnesses and synthetic data (20:03) Evaluating data when you train on reasoning traces (25:27) General versus specialized pretraining (27:08) Benchmark maxing and the GLM question (31:51) Getting diversity in, and the NVIDIA personas (35:02) Hidden reasoning traces and the fight over model IP (38:17) Mid-training and the "It's All Training" thesis (41:47) Can small models actually compete (45:01) Cybersecurity and Europe's strategic gap (47:08) Do you need a big model to orchestrate the small ones (52:08) Sovereign AI and the limits of national champions (56:42) Scaling laws when you control the data (01:00:41) The NVIDIA persona datasets (01:04:52) What you actually do with synthetic personas (01:08:22) Closing thoughts Music "Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0. About The Information Bottleneck is hosted by Ravid Shwartz-Ziv and Allen Roush, featuring in-depth conversations with leading AI researchers about the ideas shaping the future of machine learning.

    • Transcript
  • July 20 · 1 hr 7 min

    Dhruv Batra: The Browser Is a Robotics Problem - From Embodied AI at Meta to Web Agents at Yutori

    Dhruv Batra spent years leading Embodied AI at Meta, training virtual robots to navigate photorealistic 3D scans of real buildings with pure reinforcement learning. Then he left to co-found Yutori and build agents for a very different environment: the web browser. In this episode, Dhruv explains why he sees these as the same problem. Web agents, in his framing, are robots that act in a browser (pixels in, actions out), and the web turns out to be just as messy an environment as the physical world. Along the way, we cover his definition of intelligence as "navigation in idea space," why robotics is lagging LLMs, the sim-to-real gap and why you can't fake friction coefficients, the teleoperation counterexample to the "it's a sensor problem" argument, and his provocative claim that under the current paradigm, we solved machine learning and didn't even realize it. He also makes the case for why the scaling hypothesis isn't falsifiable, why JEPA-style arguments deserve to be grappled with, how Yutori trains its Navigator models with RL on live websites, and what happens to the ad-supported web when agents, not eyeballs, do the browsing. Timeline 00:01 — Intro 00:54 — What embodied AI actually means 06:47 — Intelligence as navigation in idea space 13:26 — Habitat: training robots with pure RL, no maps 20:04 — Why robotics is behind LLMs 28:24 — Sim-to-real: what you can and can't fake 33:34 — "We solved ML and nobody noticed" 37:12 — Leaving Meta, founding Yutori 43:21 — Web agents: screenshots in, actions out 48:15 — Why the web won't rebuild itself for agents 53:32 — Training Navigator: RL on live websites 1:01:04 — Who pays for the web when agents browse? 1:09:17 — What Yutori means, closing thoughts Music "Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0. About The Information Bottleneck is hosted by Ravid Shwartz-Ziv and Allen Roush, featuring in-depth conversations with leading AI researchers about the ideas shaping the future of machine learning.

    • Transcript
  • July 16 · 49 min

    How to Turn Research Into Billion-Dollar Companies, with Ion Stoica

    Ion Stoica has done what almost no academic ever does — repeatedly turned university research into billion-dollar companies. He co-founded Databricks (now valued at over $100 billion), Anyscale, Arena AI and Conviva, while his Berkeley lab produced the open source projects the entire AI industry runs on: Ray, vLLM, and SGLang. In this episode, we ask him how it's actually done. His answer is surprisingly unromantic: solve a problem people already care about, build an artifact good enough that they adopt it, and pay attention to the moment users start asking "who maintains this after the students graduate?" - that's when a project becomes a company. He's also insistent that the credit belongs to his students. From there, the conversation goes deep into what he's watching now: why the AI stack has become an order of magnitude more complex than the Hadoop/Spark era, why maximizing GPU utilization is "the name of the game" for any enterprise, and why coding agents will struggle with distributed systems long after they've mastered web apps. He shares a memorable reward-hacking story — a load balancer that maximized throughput by dropping requests — explains why the gap between open and closed models sits at about six months, and closes with his case for regulating AI by outcomes, not capabilities. Timeline 00:00 — Introduction: welcoming Ion Stoica 01:21 — The playbook: how research projects become companies 05:22 — Will vLLM and SGLang stay open source? 07:47 — The real bottleneck in the AI stack: complexity, not just hardware 14:31 — Should algorithms follow infrastructure, or the other way around? 16:13 — Can AI coding tools write distributed systems and GPU kernels? 21:09 — Verifiers, harnesses, and the limits of outsourcing understanding 25:41 — Reward hacking: the load balancer that dropped requests 25:58 — How should enterprises consume GPUs? Utilization as the name of the game 30:23 — GPU scarcity: will the compute crunch ever end? 35:27 — Hyper-optimization and the risk of locking in today's architectures 37:17 — Open vs. closed models: why every company wants to own the stack 40:35 — The six-month gap, and the rising cost of training frontier models 43:58 — Kimi, Qwen, and who's incentivized to keep open models alive 45:39 — Regulation: outcomes, not capabilities 47:41 — Self-regulation, concentration of power, and auditing open models 48:32 — Wrap-up Music "Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0. About The Information Bottleneck is hosted by Ravid Shwartz-Ziv and Allen Roush, featuring in-depth conversations with leading AI researchers about the ideas shaping the future of machine learning.

    • Transcript
  • S1 · E53
    July 13 · 59 min

    Kaggle Grandmasters, Agent Skills, and Why Everyone Is Overfitting with Jean-Francois Puget (NVIDIA)

    Jean-Francois Puget is a Director and Distinguished Engineer at NVIDIA, where he leads the Kaggle Grandmasters team, and he's ranked third on Kaggle's all-time list. We caught him on the day NVIDIA announced Nemotron Ultra and its new agent skills repo. We talk about what skills actually are, why they beat MCP tools on context cost, and how NVIDIA built an evaluation pipeline to separate skills that help from skills that don't. From there we talk about the thing JFP cares about most: evaluation. He explains why most LLM benchmarks reward overfitting, how his team discovered O3 could pick the right files to fix SWE-bench issues without reading them, and why the only benchmarks he trusts are the ones where you commit before you see the score, which is exactly how Kaggle works. He predicts a "bloodbath" for the wave of competitors letting coding agents chase leaderboard scores with no notion of validation. We also get into what coding agents are actually good for ("a mix of a genius and a dumb person"), the multi-agent system at NVIDIA that built a working PyTorch clone that runs 10x slower than the real thing, his unfiltered take on frontier lab PR and the Mythos release, whether AI is a bubble, and the story of how his team won ARC-AGI with a 4-billion-parameter model at 20 cents a task, including jumping from third to first in the final hours of a seven-month competition. Timeline 00:00 — Intro 01:05 — NVIDIA's announcements: Nemotron Ultra and the agent skills repo 07:21 — Skills vs MCP tools, and progressive disclosure 10:24 — Agents that write their own skills: a new form of learning 13:33 — When overfitting is fine (and when it isn't) 15:47 — Why most LLM benchmarks reward overfitting 17:06 — The SWE-bench contamination story: O3 picks files without reading them 19:45 — How LLMs changed Kaggle, and the coming "bloodbath" 25:40 — What makes a good data scientist: evaluation and one-bit experiments 28:56 — Running Codex at scale: the top token consumers at NVIDIA 29:37 — Did coding agents kill AutoML? 30:16 — Genius and dumb at once: the limits of coding agents 35:21 — Humans in the loop, sandboxing, and the teenage hacker who never wrote code 37:42 — Mythos, frontier lab PR, and open source 40:08 — Why NVIDIA builds open models, and where it's already frontier 43:48 — World models, robots, and the coffee test 49:20 — Why agents still can't play Dota 50:24 — Is AI a bubble? 53:14 — Winning ARC-AGI with a 4B model at 20 cents a task 57:39 — Kaggle is a legal drug Music: "Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0. About: The Information Bottleneck is hosted by Ravid Shwartz-Ziv and Allen Roush, featuring in-depth conversations with leading AI researchers about the ideas shaping the future of machine learning.

    • Transcript
  • S1 · E52
    July 9 · 1 hr 13 min

    AI Agents and The Golden Age of Asking Questions with Dimitris Papailiopoulos (MSR/UW-Madison)

    In this episode, we talked with Dimitris Papailiopoulos, researcher at Microsoft Research's AI Frontiers lab and professor at the University of Wisconsin, about doing research in the age of agents. Dimitris told us about the Sunday morning that changed how he works: he handed Claude Code and Codex a question he'd been sitting on for years, went about his day, and came back to an answer. After a few days of dread about what's left for humans, he landed somewhere more optimistic, calling this the golden age of asking questions. We talked about his "smallest transformer that can add" leaderboard, a symbolic GSM8K solver built from if-else statements, and what happened when he put two Claude Code instances in the same file system and told them to do something cool (one pair invented a communication protocol, the other played Battleship). We also got into diversity and slop in agent-generated ideas, why agents get stubborn after a million tokens, harness overfitting on Terminal-Bench, continual learning and world models, whether agents need vision, and where information theory actually helps in AI and where it's a katana used to make coffee. Timeline 00:00 Intro 01:45 How agents changed the way Dimitris does research 04:30 A Sunday morning with Claude Code, Codex, and GSM8K 07:15 The dread, then the golden age of asking questions 08:20 Taste and verification, and how we train students now 09:53 Will models make human verification obsolete? 11:30 The smallest transformer that can add 10-digit numbers 13:40 Humans as initializers for gradient descent in idea space 15:32 Allen on diversity, slop profiles, and high temperature research 21:44 When Claudes meet: Battleship, invented protocols, and a grokking paper 25:53 Single agent vs multi-agent under fixed compute 30:28 Auto-research benchmarks and what agents actually accelerate 35:14 Inside the symbolic GSM8K solver (with a live progress check) 40:04 Idea overfitting and why agents refuse to change course 44:00 Learning from failure traces and harness overfitting 48:04 Continual learning, memory files, and world models 51:30 Why don't labs personalize models on your own history? 57:52 Agent-to-agent communication: is Jira the right tool? 1:01:25 Multimodality: vision as a tool vs one unified model 1:05:40 Information theory and AI, or making coffee with a katana 1:11:23 Closing thoughts: ask bigger questions Music: "Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0. About: The Information Bottleneck is hosted by Ravid Shwartz-Ziv and Allen Roush, featuring in-depth conversations with leading AI researchers about the ideas shaping the future of machine learning.

  • S1 · E51
    July 2 · 1 hr 11 min

    Why All Models Learn the Same Thing with Phillip Isola (MIT)

    Phillip Isola, professor at MIT, joins us to talk about representation learning: what makes a representation good, why different models seem to converge on similar representations, and whether pre-training is really over. We discuss the platonic representation hypothesis and its limits, why clustering structure matters more than global geometry, and Phillip's new neural thickets paper arguing that post-training is easier than people think because pre-trained weights already sit near solutions to downstream tasks. Phillip also explains why he thinks LLMs are already world models, why he's betting on RNNs making a comeback, and why his most exciting current direction is artificial life: putting LLM agents in open environments with no fixed task and studying them like new organisms. Timeline: 00:00 Intro song 00:13 Intro 01:05 What is representation learning and why it matters 04:09 What makes a representation good: minimality and sufficiency 10:03 How cross entropy and contrastive learning shape representations 14:35 Dimensionality reduction and why dimension isn't the right complexity measure 16:35 Compression and geometric clustering during training 19:27 The platonic representation hypothesis and what actually converges 22:53 Local neighborhoods vs global structure: the Aristotelian follow-up 24:33 When convergence is strong: truth vs the space of possibility 28:09 Is there true similarity in the world? The Bouba-Kiki effect 30:56 World models vs autoregressive LLMs 32:14 Diffusion LLMs as a special case of autoregressive models 33:42 What architectures win in five years: the case for RNNs 36:11 Grad student descent, or do we actually have principles? 40:51 Feathers and wings: what to take from biology 43:17 How close are we to brain-like models? Marr's three levels 47:01 Are better models becoming less human-like? 49:38 Is pre-training all you need? The neural thickets paper 54:18 LoRA, low rank fine-tuning, and why post-training is easier than we thought 56:01 RL environments and what our benchmarks actually test 1:01:11 Artificial life: LLM agents as new organisms 1:07:20 What's overlooked in AI research right now 1:08:36 Why stay in academia, and doing science in the age of Opus Music: "Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0. About: The Information Bottleneck is hosted by Ravid Shwartz-Ziv and Allen Roush, featuring in-depth conversations with leading AI researchers about the ideas shaping the future of machine learning.

  • S1 · E50
    June 29 · 1 hr

    AI for Science with Qichao Hu (Molecular Universe / SES AI)

    Most AI-for-science companies are selling shovels. Qichao Hu wants the gold. In this episode, we talk with Qichao, the founder and CEO of Molecular Universe, the AI-for-science platform that grew out of SES AI, a high-energy-density battery developer he's run for fourteen years. His core distinction is that companies from the AI world build tools, such as foundation models that predict properties, while companies from the science world care about the final product, such as the new battery or material that actually ships. Molecular Universe sits firmly on the science side, and the difference shows up everywhere from what they publish to what they refuse to. We get into the actual workflow of materials discovery and where AI compresses it. A single trial in a traditional lab can take a year with maybe a 40% success rate; the goal is to run a thousand candidates in parallel and turn that year into a week. Qichao walks through improving low-temperature fast-charging for EV batteries: from hypothesis generation through molecule-, material-, and device-level property prediction, down to autonomous labs that synthesize and test the top candidates without a human touching a pipette. The hardest problem, it turns out, isn't predicting molecular properties or measuring device performance, but it's the black box connecting the two. In batteries, that's the solid-electrolyte interface, which the field has been hand-waving about since the seventies. And the thing standing in the way of cracking it isn't a clever training trick but data: companies sitting on twenty years of records are finding it too messy, incomplete, and poorly labeled to train on, and are having to start collecting from scratch with new protocols and robots. Timeline 00:13 — Intro and welcome; 01:19 — Shovel vs. gold 05:18 — Why the world's smartest scientist doesn't automatically give you a better battery 07:25 — The discovery workflow 09:37 — Exploration vs. exploitation 11:54 — Safety and filtering: screening novel molecules against banned and toxic-substance lists 17:55 — How hypotheses get generated, and where frontier LLMs help 20:29 — From hypothesis to ~400 formulations: property prediction, ranking, and handing off to autonomous labs 26:37 — "A foundation model for everything" — and the black box between molecular properties and device performance 30:01 — World models and physics 33:09 — The great unknown in batteries 37:08 — Simulation vs. reality: calibrating massive simulated datasets with a sliver of experimental data 41:47 — Lab robotics: how fast the hardware has caught up, and what a floor of autonomous labs looks like 43:50 — The real bottlenecks 50:21 — Pre-training from scratch vs. post-training LLMs, and why training tricks haven't reduced the need for good data 52:42 — Evaluation 55:42 — Publish the B+ model, keep the A model 58:05 — Five years out 1:00:37 — Closing thoughts and wrap Music: "Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0. About: The Information Bottleneck is hosted by Ravid Shwartz-Ziv and Allen Roush, featuring in-depth conversations with leading AI researchers about the ideas shaping the future of machine learning.

  • S1 · E49
    June 24 · 1 hr 5 min

    Infrastructure for AI at Scale - With Benny Chen (Fireworks AI)

    We talk a lot on this show about RL, agents, and the move between pre-training and post-training, but not enough about the layer everything actually runs on. Benny Chen, co-founder of Fireworks AI, one of the largest inference platforms around, walks us through what it takes to serve models at scale: sourcing GPUs, writing the kernels, the runtime, and the routing layer that lets a customer hit one endpoint and forget the rest. We talk why the real bottleneck is power, not chips, and why that favors Nvidia and Google. Why MoE keeps winning even when dense models look better on paper and why he'd rather run fungible capacity at 95% than specialized chips at 60%. We also talk about quantization limits, where RL efficiency has to go next, and his case that AI is still under-hyped. We also get into cross-region training, sparse autoencoders and why interpretability hasn't taken off in open source, whether open models can close the gap, and a frank read on Anthropic's go-to-market. Timeline 00:00 — Intro: the part of AI nobody talks about 01:20 — What "infrastructure for AI" actually means: the layers, from GPUs up to routing 02:59 — Why not just buy your own GPUs and do it yourself? 05:17 — The scale Fireworks runs at 06:35 — Hardware inflation, GPU costs, and the real risk hiding in commit duration 10:14 — Nvidia vs AMD vs TPUs, and why power is the bottleneck 11:57 — Mixing GPU types and generations; fungibility vs. specialization 14:22 — Once you have the GPUs, what's the next layer to build? 17:04 — Dense vs. MoE, and why the hardware picks the winner 21:07 — Quantization: is FP4 the floor? TurboQuant and INT vs. FP 24:28 — How tied are the algorithms to the hardware? 25:12 — DeepSeek, DeepGEMM, and next-token prediction as reconstruction loss 28:50 — Why RL is still wildly inefficient compared to pre-training 30:08 — Speculative decoding, AI-generated kernels, and auto-research 34:00 — The AGI question: why text gets automated but vision may stay expensive 37:07 — Hype check: why Benny thinks AI is still under-hyped 41:28 — Training vs. inference at the infrastructure level 44:12 — Scaling across data centers: cross-region training with Cursor 45:40 — Sparse autoencoders, interpretability, and why open source is human-constrained 49:04 — Will open models catch up — on quality and on compute? 51:41 — Are we plateauing? Opus 4.7 vs. 4.6 and the coming data wars 54:41 — Physical limits, HBM, and whether chips keep getting faster 58:17 — The belief about inference everyone gets wrong 59:31 — Anthropic, mythos, and a frank take on go-to-market 1:04:41 — Wrap-up Music: "Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0. About: The Information Bottleneck is hosted by Ravid Shwartz-Ziv and Allen Roush, featuring in-depth conversations with leading AI researchers about the ideas shaping the future of machine learning.

  • S1 · E48
    June 21 · 1 hr 18 min

    Broken Peer Review, AI, and Worms — with Oded Rechavi

    Oded Rechavi is a biologist at Tel Aviv University and the co-founder of QED, a company building AI to review scientific work. He's also spent years studying worms. We start with what's wrong with peer review and grant funding: why it takes years to publish, why reviewers are often your own competitors, and why the whole thing is locked to an economic model that rewards publishing more papers, not better ones. Oded explains why he doesn't call QED "peer review" at all, and what it would take to actually validate science instead of just stamping it. Then we get into the biology. C. elegans has exactly 959 cells, every one of them named, and a fully mapped brain. Oded's lab studies how a worm's experiences get passed to its offspring through RNA rather than DNA — meaning what happens to a worm in its lifetime can change its descendants. We also talk about using ancient DNA to reassemble the Dead Sea Scrolls, what AI can and can't do for biology, and why he wants to build an "Ironman suit" for researchers rather than replace them. 00:00 Intro 01:35 Why scientific publishing is broken 04:02 Years to publish, and what it costs science 07:20 Bad reviewers, conflicts of interest, and the money 10:47 Why preprints don't fix it 15:37 How AI conferences handle review 22:07 Conferences vs. journals — does slow review help? 25:22 Building QED: review, not peer review 30:02 Tracking a paper from idea to submission 33:11 What writing a grant actually involves 35:00 The ERC reviewer crisis 37:06 Tailoring feedback to your field 41:48 Switching to biology 44:30 Every cell has a name: inside C. elegans 46:28 Inheritance without DNA 48:16 What the worm "thinks" changes its offspring 51:58 Reassembling the Dead Sea Scrolls with ancient DNA 56:07 Psychedelics and worms 58:36 Can AI run the research itself? 1:04:49 Automation vs. validation 1:07:12 The origin of life 1:08:49 Why people reject AI-written work 1:16:18 Will humans still have a role? 1:17:39 Wrap-up Music: "Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0. About: The Information Bottleneck is hosted by Ravid Shwartz-Ziv and Allen Roush, featuring in-depth conversations with leading AI researchers about the ideas shaping the future of machine learning.

  • S1 · E47
    June 16 · 1 hr 29 min

    Will AI Take Our Jobs? With Alex Imas (Google/University of Chicago)

    Will AI take our jobs? We put the question to Alex Imas, the new Director of AGI Economics at Google DeepMind and a professor at Chicago Booth, whose entire job now is studying how frontier AI reshapes the economy. His short answer: probably some of them, but the popular story is mostly wrong about which jobs and how fast. Alex makes the case that a job is a bundle of tasks, not a single thing AI either does or doesn't do, and that the number of people who should actually care about is how much consumer demand responds to falling prices. Get that wrong and you predict mass layoffs. Get it right and you sometimes predict more hiring. We get into why the automation panic is two centuries old, why he thinks blue-collar work is in more danger than white-collar, and why the people already winning are the ones adopting AI fastest. We also cover the AGI versus ASI distinction and why it changes everything for the economy, what happens when there's no moat and open models stay six to eight months behind, the three-tier pricing future he sees coming after the 2026 compute crunch, and what any of this means if you're deciding whether to send your kids to college. The episode was recorded before Alex joined Google Timestamps 00:00 Meeting Alex Imas 00:44 Will AI take our jobs? 03:35 Is this an AI question or an economics question? 06:18 The economy is already behind the AI we have 07:43 Why AI adoption is K-shaped 12:51 Was Andrew Yang right? 13:45 The automation panic is 200 years old 16:46 Dario's six-month claim, and why we don't see it yet 17:22 A job is not a task 22:38 The three numbers that actually predict the labor market 22:42 The chess engine analogy and the centaur phase 25:45 Recursive self-improvement and the hamburger problem 30:06 Should AI labs be the ones answering alignment questions? 31:17 The "invisible hand wave" and why nobody wants fully autonomous AI 33:27 AGI vs ASI, and why the difference is everything 35:28 Commodities vs relational goods 41:14 Star Trek, replicators, and predicting with sci-fi 45:20 Inequality and the Upper West Side VCs 46:21 Your money manager was automated in the 1960s 50:47 Are OpenAI and Anthropic overvalued? The moat problem 54:29 What has to be true for the losses to make sense 55:43 Cognitive atrophy and monopoly fears 57:00 The 2026 compute crunch and the three-tier pricing future 1:01:52 The Apple vs Android analogy 1:03:54 A rich-country perspective 1:04:16 Protecting the skills that actually matter 1:07:02 Will not using AI become a status symbol? 1:08:53 Does capitalism even survive? 1:13:44 Redistribution becomes the political battleground 1:18:16 Blue collar vs white collar: who's really at risk 1:21:18 Advice for parents in an AI world 1:22:43 Saving for retirement when the Valley says don't 1:25:06 Will non-elite colleges survive? Music: "Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0. About: The Information Bottleneck is hosted by Ravid Shwartz-Ziv and Allen Roush, featuring in-depth conversations with leading AI researchers about the ideas shaping the future of machine learning.

  • S1 · E46
    June 13 · 1 hr 19 min

    Why AI Benchmarks Are Lying to You - with Wenhu Chen (Meta/University of Waterloo)

    In this episode, we sit down with Wenhu Chen, research scientist at Meta MSL, assistant professor at the University of Waterloo, and the person behind MMLU-Pro and MMMU. If you've read a frontier model release in the last two years, you've seen his benchmarks. That makes him one of the best people to answer the question everyone dances around: when a model jumps from 40% to 90% on your benchmark, how much of that is real? In this episode, we dig into why benchmarks have become the loss function of the entire field - design a bad one, and thousands of brilliant researchers will spend months hill-climbing in the wrong direction. Wenhu is surprisingly candid about the limits of his own creations: contamination is everywhere, saturation turns frontier benchmarks into unit tests, and popular alternatives, such as LM Arena, mostly measure tone and length rather than capability. His answer is to evaluate models where they've never been: private codebases, hospital data, and the messy, live internet. We also talk about ClawBench, his new benchmark that deploys agents to over 140 real production websites to do things people actually want done, such, such as ordering food, booking tickets, and applying for jobs. The best model in the world completes about a third of these tasks. We unpack why: bot detection, models that refuse to click "pay," agents that give up the moment an environment doesn't match their training, and harnesses that can swing results by 20% without changing the model at all. Along the way, we cover the overlooked science of evaluating pre-training, data flywheels, and synthetic environments for agent training, and whether RL teaches models to reason or just surfaces what's already there. We close with Wenhu's predictions: exploration and adaptability will improve rapidly, but security will become the field's hardest problem as agents gain real permissions in the real world. Timestamps 00:00 – Intro 00:55 – What good evaluation means, and how it's changed since the early GPT days 03:35 – Benchmarks as the field's loss function 05:50 – Contamination: the problem nobody fully solves 08:08 – MMLU-Pro scores: real progress or training on the test set? 11:05 – Can you measure creativity? 12:34 – Why human judges and arenas are unreliable — and what to use instead 19:22 – What a good benchmark actually looks like 22:34 – Chain of thought: signal or scratchpad? 26:01 – Auto-research and hill-climbing agents 28:52 – Harnesses: 20% swings without touching the model 32:28 – Safety, model release, and an "FDA for models" 36:53 – The overlooked science of pre-training evaluation 43:49 – Designing pre-training benchmarks when one run costs a billion dollars 49:45 – ClawBench: agents on 140+ live websites, and why the best model gets 33% 54:42 – How MMLU-Pro and MMMU-Pro were born from public complaints 59:16 – Pixel agents vs. APIs: will MCP kill computer use? 1:02:11 – Training agents: data flywheels and synthetic environments 1:05:43 – SFT vs. RL, and does RL teach reasoning or reveal it? 1:09:21 – What gets solved next year — and what doesn't 1:14:32 – Undervalued ideas, and what's next for ClawBench Music: "Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0. About: The Information Bottleneck is hosted by Ravid Shwartz-Ziv and Allen Roush, featuring in-depth conversations with leading AI researchers about the ideas shaping the future of machine learning.

  • S1 · E44
    June 7 · 1 hr 29 min

    Jürgen Schmidhuber - Part 2: JEPA, the Road to AGI, and Who Really Invented Modern AI

    In the second half of our conversation with Jürgen Schmidhuber, we focus on the key ideas he's pursued since the early 1990s and discuss why he believes these concepts are only now being rediscovered. We start with JEPA. Jürgen argues that the method LeCun named in 2022 is the same family he published in 1992 as Predictability Maximization. From there he traces the adversarial lineage back further still, to his 1990 world-model paper and 1991 Predictability Minimization - the curiosity-driven minimax games he sees as the real origins of GANs. We also talk about why these ideas took thirty years to land, why today's trillion-dollar data-center buildout is driven by AGI fear, and why he thinks Apple may come out ahead. The back half turns to what he sees as the real frontier: physical AI. Today's systems are superhuman behind the screen but helpless at a leaky pipe, and until a robot can use human tools, there's no AGI. He discusses self-replicating, self-improving machines as "a new kind of life," reframes continual learning and test-time training as ideas from his 1991 fast-weight work, and detours through Solomonoff's universal prior, Hutter's AIXI, and the Gödel machine. We close on the subject Jürgen is famous for: scientific credit. He makes his case for rigorous attribution, casts himself as a "speaker for the dead" championing forgotten pioneers like Ivakhnenko, and reflects candidly on whether the fights are personal. Timeline 00:30 — What JEPA is, and the 1992 Predictability Maximization story 04:54 — Implementing PMAX: autoencoders, Siamese networks, Infomax 09:10 — Predictability Minimization, factorial codes, and the roots of GANs 16:00 — Why it took 30 years: the economics of compute 20:52 — Data, the web, and 1990 as the origin point 23:09 — Hardware inflation, the trillion-dollar buildout, and the coming crash 34:05 — Physical AI: the plumber problem and self-replicating machines 41:14 — Which 90s ideas are being scaled right now 45:26 — Continual learning and test-time training as "old hats" 55:19 — Measuring intelligence: Solomonoff, AIXI, and the Gödel machine 1:05:26 — Self-replication and von Neumann 1:09:51 — Will he see AGI in his lifetime? 1:10:42 — Credit, integrity, and being a "speaker for the dead" Music: "Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0. "Palms Down" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0. Changes: trimmed About: The Information Bottleneck is hosted by Ravid Shwartz-Ziv and Allen Roush, featuring in-depth conversations with leading AI researchers about the ideas shaping the future of machine learning.

    • Transcript
Showing 1–20 of 24 episodes