Skip to content
Artwork for Embodied AI 101
Embodied AI 101 · Monday · 29 min

BridgeVLA++: Keep the Geometry, Give the Robot a Memory

Renders point clouds as three orthographic 2D views fed into a pretrained PaliGemma VLM, predicting heatmaps to recover 6-DoF actions with temporal/spatial memory modules. Achieves 95.4% success on 13 real Franka tasks with only 3 demos per task, outperforming π0, RVT-2, and 3D Diffuser Actor, with fully open weights and code.

0:00-29:54

transcript

No transcript — this publisher did not publish one.

show notes

Renders point clouds as three orthographic 2D views fed into a pretrained PaliGemma VLM, predicting heatmaps to recover 6-DoF actions with temporal/spatial memory modules. Achieves 95.4% success on 13 real Franka tasks with only 3 demos per task, outperforming π0, RVT-2, and 3D Diffuser Actor, with fully open weights and code.