Latent Action Pretraining from Videos (ICLR 2025): the first unsupervised method for pretraining vision-language-action models from action-less video — a VQ-VAE learns discrete latent actions between frames, a VLM (7B backbone) is pretrained to predict them, then a small robot dataset grounds them in real actions, with ~30x better pretraining efficiency than conventional VLA pretraining. Won the CoRL 2024 LangRob workshop Best Paper, and the latent-action idea was adopted in NVIDIA GR00T N1. KAIST LK Lab lead (Minjoon Seo) with Microsoft Research, NVIDIA, and UW.

Paper

Venue ICLR 2025
roboticsmultimodalopen-weight

Related