Skip to content
Artwork for Embodied AI 101
Embodied AI 101 · Thursday · 29 min

THAW-VLA: World-Model Features Without World-Model Latency

This work distills frozen world-model representations into compact VLAs via a single feature-alignment loss at training time, with the teacher cached once and absent at inference; a 0.8B student achieves 97.9% on LIBERO and lifts RoboCasa-GR1 humanoid performance from 48.2% to 50.5% at zero extra runtime cost (32 ms / 1.86 GB on RTX 5090). The approach generalizes across scales and architectures on both single-arm and bimanual real hardware.

0:00-29:21

transcript

No transcript — this publisher did not publish one.

show notes

This work distills frozen world-model representations into compact VLAs via a single feature-alignment loss at training time, with the teacher cached once and absent at inference; a 0.8B student achieves 97.9% on LIBERO and lifts RoboCasa-GR1 humanoid performance from 48.2% to 50.5% at zero extra runtime cost (32 ms / 1.86 GB on RTX 5090). The approach generalizes across scales and architectures on both single-arm and bimanual real hardware.