
Embodied AI 101 · Yesterday · 31 min
Stop Teaching the VLA Your Camera: Robot-Centric Pointmaps as an Action-Aligned Visual Interface
0:00-31:22
transcript
show notes
Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in the robot's own 3D coordinate frame, yet most VLAs observe the scene in 2D. This paper proposes robot-centric pointmaps to bridge this gap for vision-language-action models.