Skip to content
Artwork for Embodied AI 101
Embodied AI 101 · August 31 · 30 min

SIRModel and the Missing Geometry Layer in Vision–Language Manipulation

Long-horizon robotic manipulation requires a policy to bridge task-level semantic reasoning with metric three-dimensional interaction geometry. Existing vision–language–action policies usually acquire geometry implicitly from visual tokens or introduce deterministic intermediate representations that lack spatial grounding. This paper proposes SIRModel, which learns a spatial intermediate representation to parameter-efficiently fine-tune a vision language model for robotic manipulation tasks.

0:00-30:30

transcript

No transcript — this publisher did not publish one.

show notes

Long-horizon robotic manipulation requires a policy to bridge task-level semantic reasoning with metric three-dimensional interaction geometry. Existing vision–language–action policies usually acquire geometry implicitly from visual tokens or introduce deterministic intermediate representations that lack spatial grounding. This paper proposes SIRModel, which learns a spatial intermediate representation to parameter-efficiently fine-tune a vision language model for robotic manipulation tasks.