RoboTok retrieves human hand-pose trajectories from internet videos to automatically generate scalable robot training data for dexterous manipulation tasks. This approach enables large-scale data collection without manual annotation or teleoperation.
VLAct proposes a representation-centric continued pre-training approach for Vision-Language-Action models that converts limited robot data into transferable action knowledge, achieving 92.5% success on RoboTwin with only 20% of full training data. The met
Breaks intractable full world-model RL into a heavy global trajectory model combined with a lightweight latent local dynamics approximator, enabling scalable RL for contact-rich humanoid skills without backpropagating through the full model.
Pantheon Industries introduces Reinforced Planning (RP-1), the first fully learned planner that iteratively improves robot action plans from scratch using Reinforcement Learning over a frozen World Model. RP-1 reimagines planning as a learned policy: star
Vision–Language–Action (VLA) policies remain brittle under modest distribution shift. On LIBERO-Plus, contemporary models that solve clean tasks at high rates can fall below 30% success when the camera's viewpoint or background changes. This paper propose
Introduction Learning robust visuomotor policies for bimanual manipulation remains challenging due to the stringent requirements for precise coordination between arms and the ability to generalize across diverse environmental conditions. Existing approach
Long-horizon robotic manipulation requires a policy to bridge task-level semantic reasoning with metric three-dimensional interaction geometry. Existing vision–language–action policies usually acquire geometry implicitly from visual tokens or introduce de
Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks. Most hierarchical vision-language-action (VLA) models make each such decision with a fixed compute budget. τ0-VLA
Introduces dynamic Roll/Stop decision-making in world action models that adapts imagination based on planning benefit, risk, and compute cost for more efficient robot planning. Enables world action models to selectively engage imagination only when benefi
Learning sensorimotor contingencies—that is, the link between one's actions and their sensory effects—is fundamental to developing body knowledge, understanding causality, and developing a sense of agency. In developmental psychology, this...
High-voltage grid maintenance and chemical processing demand dexterous manipulation in confined, visually occluded environments. Existing teleoperation systems present a severe cost–fidelity trade-off: industrial platforms exceed USD 60,000 yet lack spati
Search-and-rescue (SAR) operations in disaster environments require drone swarms to coordinate efficiently despite incomplete information and potential communication failures. Existing stigmergy-based approaches provide low-bandwidth coordination but rely