Skip to content
Artwork for RobTalk
RobTalk · Thursday · 45 min

World Models & the next step for Physical AI

Is 2026 the year world models make Physical AI real? Vision-Language-Action models have become the standard approach to robot learning over the last few years. In this episode, Felix Frank from our Robot Intelligence team explains why the next shift is already underway: world models that predict what happens next before a robot decides what to do. We cover how the field moved from small, task-specific models to large pretrained architectures, why memory and context work differently in robotics than in language models, and why robots still need to run at high frequency with very limited context per decision. We also break down the two competing approaches to building a world model: one that generates full video in pixel space, and one that predicts directly in a compressed latent space, closer to the JEPA approach. Finally, we talk about where world models actually help today: generating synthetic training data, predicting the action needed to reach a desired future state, and simulating multiple possible outcomes before choosing one to execute. You'll gain insights into: Why Vision-Language-Action models became the default approach in robotics How robots handle memory and context differently than language models Why control frequency, not context length, is the real constraint in robotics Two different ways to build a world model, and the trade-offs between them Why touch and force sensing matter as much as vision for some tasks Why mechanical and electrical engineering remain the next big bottleneck More about RobCo: Website: https://www.rob.co LinkedIn: https://www.linkedin.com/company/robco-therobotcompany/ Instagram: https://www.instagram.com/robco_therobotcompany/ physicalAI #robco #robotics #autonomy #podcast #worldmodels 01:38 The 2026 state of Physical AI 02:50 From task-specific models to the ChatGPT moment 07:09 The data bootstrapping problem in robotics 09:21 What goes into a Vision-Language-Action model 11:15 Why "state" matters, and the Markov assumption 17:10 Context length: robotics vs. language models 18:43 Inside the control hierarchy: motors to orchestration 20:05 Onboard GPUs vs. cloud compute 20:44 The human latency baseline 21:11 What is a world model, really? 22:34 Video generation as a world model 25:30 Is a language model already a world model? 27:02 Two ways to build a world model 30:48 Three ways world models help robots today 34:08 What happens in the next six to twelve months 35:35 Handling multiple possible futures 37:52 Beyond vision: touch and force sensing 39:19 The real bottleneck: mechanical engineering 40:46 How fast the field moves, and RobCo's 24/7 goal 42:36 Inside RobCo's experimentation framework

0:00-45:44

transcript

No transcript — this publisher did not publish one.

show notes

Is 2026 the year world models make Physical AI real? Vision-Language-Action models have become the standard approach to robot learning over the last few years. In this episode, Felix Frank from our Robot Intelligence team explains why the next shift is already underway: world models that predict what happens next before a robot decides what to do. We cover how the field moved from small, task-specific models to large pretrained architectures, why memory and context work differently in robotics than in language models, and why robots still need to run at high frequency with very limited context per decision. We also break down the two competing approaches to building a world model: one that generates full video in pixel space, and one that predicts directly in a compressed latent space, closer to the JEPA approach. Finally, we talk about where world models actually help today: generating synthetic training data, predicting the action needed to reach a desired future state, and simulating multiple possible outcomes before choosing one to execute.

You'll gain insights into:

  • Why Vision-Language-Action models became the default approach in robotics
  • How robots handle memory and context differently than language models
  • Why control frequency, not context length, is the real constraint in robotics
  • Two different ways to build a world model, and the trade-offs between them
  • Why touch and force sensing matter as much as vision for some tasks
  • Why mechanical and electrical engineering remain the next big bottleneck

More about RobCo: Website: https://www.rob.co LinkedIn: https://www.linkedin.com/company/robco-therobotcompany/ Instagram: https://www.instagram.com/robco_therobotcompany/

physicalAI #robco #robotics #autonomy #podcast #worldmodels

01:38 The 2026 state of Physical AI 02:50 From task-specific models to the ChatGPT moment 07:09 The data bootstrapping problem in robotics 09:21 What goes into a Vision-Language-Action model 11:15 Why "state" matters, and the Markov assumption 17:10 Context length: robotics vs. language models 18:43 Inside the control hierarchy: motors to orchestration 20:05 Onboard GPUs vs. cloud compute 20:44 The human latency baseline 21:11 What is a world model, really? 22:34 Video generation as a world model 25:30 Is a language model already a world model? 27:02 Two ways to build a world model 30:48 Three ways world models help robots today 34:08 What happens in the next six to twelve months 35:35 Handling multiple possible futures 37:52 Beyond vision: touch and force sensing 39:19 The real bottleneck: mechanical engineering 40:46 How fast the field moves, and RobCo's 24/7 goal 42:36 Inside RobCo's experimentation framework

links3