Skip to content
Artwork for Embodied AI 101
TechnologyScience

Embodied AI 101

Shaoqing Tan

Stay in the loop on research in AI and physical intelligence.

Play
  • 44 episodes
  • Avg 29 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • Yesterday · 12 min

    The Hard Part Was the Stack: PI's Robot Goes to Work

    Physical Intelligence's mobile π robot runs fully autonomous for hours in a real production environment stacking boxes at Dandelion Chocolate, exposing generalization and reliability gaps between lab demos and production utility. Represents a significant

    • Transcript
  • Yesterday · 13 min

    ZDTaichu5.0-9B: Better Spatial Reasoning, an Unfinished Edge Story

    ZDTaichu5.0-9B is a 9B edge-deployable multimodal model built on a Qwen3.5 backbone with C-RADIOv4 vision encoder that leads open sub-10B VLMs on spatial-reasoning benchmarks including ViewSpatial and MMSI-Bench, while supporting embodied AI and tool-use

    • Transcript
  • Tuesday · 28 min

    Light-Loco-Parkour: Learning the Skill Is Only Half the Problem

    Light-Loco-Parkour presents a multi-skill distillation approach for learning perceptive whole-body parkour and locomotion skills deployable on real robots, enabling versatile agile movement across diverse terrains. The method advances embodied locomotion

    • Transcript
  • Tuesday · 29 min

    OM-1: Human Skills, Many Bodies

    OM-1 is a robot foundation model trained exclusively on human manipulation data (no teleoperation or robot demonstrations) that achieves near human-level dexterity and zero-shot generalization across tabletop arms, industrial arms, and humanoids, includin

    • Transcript
  • Monday · 12 min

    Show-Harness Gives VLMs a Robot Keyboard

    Uses discrete semantic actions paired with embodiment-specific interpreters to turn any pretrained VLM into a robot controller without a dedicated policy network or additional pretraining. Demonstrates strong zero-shot generalization across diverse tasks

    • Transcript
  • Monday · 29 min

    BridgeVLA++: Keep the Geometry, Give the Robot a Memory

    Renders point clouds as three orthographic 2D views fed into a pretrained PaliGemma VLM, predicting heatmaps to recover 6-DoF actions with temporal/spatial memory modules. Achieves 95.4% success on 13 real Franka tasks with only 3 demos per task, outperfo

    • Transcript
  • Sunday · 28 min

    The Sim-to-Real Gap Inside the Gearbox

    Reinforcement learning (RL) has become a powerful tool for quadrupedal locomotion, and a sim-to-real approach is widely adopted to avoid hardware damage during training. However, the sim-to-real gap remains a significant challenge.

    • Transcript
  • Sunday · 28 min

    A Stable Grasp Isn't Enough: Giving 6-DoF Grasping a Purpose

    Robotic grasping in cluttered scenes is dominated by methods that optimise either stability or a predefined task category, leaving a gap between 'can this grasp hold the object' and 'can this grasp serve the task'. This paper bridges that gap using founda

    • Transcript
  • Sunday · 29 min

    HiDream-O1-Embodied: A World Model Must Earn Its Actions

    HiDream-O1-Embodied is a unified native architecture that ingests text, images, and video and directly outputs physical actions, positioning itself as a step beyond passive world modeling toward active embodied interaction. The architecture aims to close

    • Transcript
  • Sunday · 13 min

    Axis's Real2Sim2Real Bet: Make the Data Prove Itself

    Axis introduced a Real2Sim2Real pipeline that converts real objects and scenes into high-fidelity physics-accurate simulations and validates learned policies back on hardware, paired with a performance-driven active trajectory curation engine. This data f

    • Transcript
  • Saturday · 30 min

    A Rubik's Cube, Two Robot Hands, and One Category Error

    GPT-6 Astra is demonstrated solving a Rubik's Cube using dexterous robot hands in a real physical setup, showcasing a significant leap in vision-language-action integration for fine-grained dexterous manipulation. This demo highlights the frontier of embo

    • Transcript
  • Saturday · 33 min

    UnifoLM: Predict the Interaction, Not the Whole World

    Unitree's UnifoLM is a 6B-parameter world model trained on 2,500 hours of real robot data that predicts only dynamic pixels and optical flow rather than full frames, enabling efficient robot action prediction. This interaction-centric design reduces compu

    • Transcript
  • Friday · 32 min

    Stop Making the Video Model Remember the World

    A new world model architecture that decouples world-state evolution from rendering: an agent writes executable rules, a persistent engine tracks entities, and a video model renders the output. This separation enables more controllable and interpretable wo

    • Transcript
  • Friday · 13 min

    Atlas Turns Phone Scans Into Robot Worlds—But Not Yet Into Physics

    Atlas is a unified spatial intelligence model that reconstructs real scenes from 2–3 images, generates 1440p video along camera paths, and supports real-to-sim robotics workflows using Gaussian Splats to create photorealistic 3D digital twins from phone s

    • Transcript
  • September 10 · 33 min

    W²G-Net: Carrying Instance Identity from Shiny Scrap to the Gripper

    Purpose This study aims to develop a robust vision-guided framework for instance-level recognition and grasp-based classification of metal waste in dense, multi-scale and highly reflective industrial environments, addressing the challenges...

    • Transcript
  • September 10 · 33 min

    When Avoidance Is Not Enough: T-CARE Gives UAV Swarms a Right-of-Way

    Multi-agent UAV navigation in cluttered environments is challenged by deadlock, starvation-like persistent yielding, oscillation, and unsafe congestion, particularly in narrow passages and dynamically emerging bottlenecks. Existing approaches often rely o

    • Transcript
  • September 9 · 31 min

    Teaching Robots What "Better" Means

    A method for preference-based learning applied to dexterous robotic manipulation, enabling robots to learn from human preferences in freeform settings. Discussed publicly around September 8, 2026.

    • Transcript
  • September 9 · 29 min

    TANGO: Stop Treating a Humanoid Like a Disk on the Floor

    TANGO directly predicts 29-DoF joint actions from RGB observations for zero-shot real-world humanoid navigation, enabling coordinated arms, torso, and gait control in cluttered 3D spaces. Presented at CoRL 2026.

    • Transcript
  • September 8 · 34 min

    One Policy, Many Bodies: Qwen-VLA's Bid to Unify Embodied AI

    Embodied intelligence is often studied through specialized models for individual tasks such as manipulation or navigation, resulting in fragmented capabilities and limited generalization across tasks, environments, and robot embodiments.

    • Transcript
Showing 1–20 of 44 episodes