Skip to content
Artwork for Embodied AI 101
Embodied AI 101 · Yesterday · 31 min

Stop Teaching the VLA Your Camera: Robot-Centric Pointmaps as an Action-Aligned Visual Interface

Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in the robot's own 3D coordinate frame, yet most VLAs observe the scene in 2D. This paper proposes robot-centric pointmaps to bridge this gap for vision-language-action models.

0:00-31:22

transcript

No transcript — this publisher did not publish one.

show notes

Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in the robot's own 3D coordinate frame, yet most VLAs observe the scene in 2D. This paper proposes robot-centric pointmaps to bridge this gap for vision-language-action models.