LAC-WM

ICML 2026

Cross-Embodiment Robot Foundation
World Models with Latent Actions

LAC-WM learns a unified latent action space from human and robot data for cross-embodiment world modeling and planning.

Huang Huang · Sriram Yenamandra · Arjun Majumdar · Elie Aljalbout · Tushar Nagarajan · Tsung-Yen Yang · Akshara Rai · Michael Rabbat · Li Fei-Fei · Jiajun Wu · Tingfan Wu · Franziska Meier

Stanford University · Work partially done at Meta FAIR Robotics

Abstract

Robot embodiments have different action spaces, making it difficult to train world models on heterogeneous data. We introduce the Latent Action-Conditioned Robot World Model (LAC-WM), which learns a shared action representation from visual transitions with auxiliary motion supervision. After pretraining on EgoDex, Agibot, and Droid, an action projector adapts the model to an unseen robot embodiment. Compared with explicit action conditioning (EAC-WM), LAC-WM improves downstream planning performance on dexterous manipulation and modified LIBERO tasks. Its performance also improves as the number of pretraining embodiments increases.

Method

LAC-WM combines an inverse dynamics model, a forward dynamics model, a motion decoder, and an action projector.

LAC-WM architecture: an inverse dynamics model extracts latent actions from observations; a motion decoder provides supervision; a forward dynamics model predicts future observations; an action projector adapts unseen robot actions.
LAC-WM architecture. Pretraining learns latent actions from observations and motion labels. Finetuning maps actions from an unseen embodiment into the pretrained latent space.

Inverse dynamics model

An inverse dynamics model encodes consecutive observations into a continuous latent action.

Motion supervision

A motion decoder reconstructs motion labels, while cross-augmentation discourages visual shortcuts.

Action projector

An action projector maps raw actions from an unseen robot into the shared space for imagined rollouts.

Evaluation

The world model ranks candidate actions sampled by a vision-language-action policy.

LIBERO

Unseen objects & locations

+9.6 points

Average task success (%)

LIBERO average task success rateY-axis 74 to 100 percent. VLA|random: 74.4 percent, VLA|mean: 81.6 percent, EAC-WM: 82.4 percent, LAC-WM: 92 percent7480859095100VLA|random: 74.4%74.4%VLArandomVLA|mean: 81.6%81.6%VLAmeanEAC-WM: 82.4%82.4%EAC-WMLAC-WM: 92%92%LAC-WM

125 episodes across five unseen task variants. All methods use the same π₀.₅ backbone. Axis: 74–100%.

Dexterous manipulation

Unseen instances & categories

+46.7% relative

Average task success (%)

Dexterous average task success rateY-axis 10 to 25 percent. VLA|random: 15 percent, VLA|mean: 17 percent, EAC-WM-S: 18 percent, EAC-WM: 15 percent, LAC-WM: 22 percent10152025VLA|random: 15%15%VLArandomVLA|mean: 17%17%VLAmeanEAC-WM-S: 18%18%EAC-WM-SEAC-WM: 15%15%EAC-WMLAC-WM: 22%22%LAC-WM

Bimanual Franka with Allegro hands. Table 2 means across three seeds, averaged over both splits. Axis: 10–25%.

Values reproduced from Tables 2 and 4. Both bar charts use truncated y-axes; labels show the reported absolute success rates. LIBERO gain is measured against EAC-WM.

LIBERO results by task
Success rate (%) · 25 rollouts per task
TaskVLA-randomVLA-meanEAC-WMLAC-WM
Book76768092
White bowl76929296
Wine bottle72848496
Alphabet soup80728088
Salad dressing68847688
Average74.481.682.492.0

Latent action representations

Explicit action encoders separate embeddings by dataset. LAC-WM’s inverse dynamics model, trained with motion decoding, shows stronger alignment across Droid, Agibot, and EgoDex.

The learned latent actions can be transferred between humans and robots. In the scaling experiments, LAC-WM improves with additional pretraining embodiments, whereas EAC-WM performance decreases.

DroidAgibotEgoDex
Four UMAP plots comparing action representations. Explicit encoders produce separated dataset clusters; the inverse dynamics model with motion decoding yields overlapping clusters.
UMAP of 7,000 action embeddings from Droid, Agibot, and EgoDex.

Qualitative results