WORLD ACTION MODELS · 2026

LD4WAM

Learning Latent Dynamics from Human Videos
for World Action Models

Zhenhao Shen1*, Jiaqi Liang1,2*, Jasper Lu1*, Feng Jiang1*, Yuran Wang1,2, Chuanbo Wei1, Jiayi Liu1, Jianchun Yang5, Qize Yu1, Jiadi You6,
Ce Hao4, Guanqi He2, Chen Xie3, Ruihai Wu1†
1Peking University2WUJI3Lightwheel4Beijing Zhongguancun Academy5Wuhan University6The University of Hong Kong

* Equal contribution.   † Corresponding author.

THE MOTION-ALIGNED BRIDGE
LDMlearn the bridge
INPUTHuman + robot
video
appearance-rich observations
semantic reconstruction
+ motion alignment
LEARNED REPRESENTATIONLatent
dynamics
semantic change + Delta EE
offline latent-dynamics targets
WDAMuse the bridge
VIDEO EXPERTGenerate
future video
pretrained visual prior
LATENT DYNAMICS EXPERT
learnable query
extract latent dynamics
ACTION EXPERTGenerate
robot action
future video + latent dynamics
01 / THE GAP

Human videos give us scale.
Robot data gives us action.

LD4WAM learns what changes between frames while suppressing appearance and embodiment details, turning diverse video into a compact signal that a robot can use.

01

Human video

Diverse and inexpensive, but pixel prediction is not directly actionable.

scale
02

Robot action

Executable and precise, but expensive to collect and tied to one embodiment.

control
02 / THE BRIDGE

Motion-aligned
latent dynamics.

Not pixels. Not raw actions.
A representation of meaningful change.

THE CORE REPRESENTATION

What happens next?

Semantic reconstruction captures high-level evolution. Motion alignment grounds it in real end-effector movement. Together they form an embodiment-agnostic bridge.

semantic changemotion alignment
zt
zt+n
zt+2n
semantic space
motion
03 / OVERVIEW
Overview of LD4WAM: curated human and robot data, latent dynamics model, and performance results
04 / ARCHITECTURE

Three experts.
One shared direction.

Learnable queries carry motion-aligned dynamics from generated futures into the action expert, while preserving the full video prior.

01VIDEO EXPERT

Imagine

Predict future observations and retain visual priors.

shared self-attention
02LATENT DYNAMICS EXPERT

Distill

Learnable queries summarize what will happen next.

information bridge
03ACTION EXPERT

Act

Use generated future video and latent dynamics to condition robot action generation.

TRAINING PROCEDURE
Stage IPre-training · all human + robot data
Stage IIAlign training · general robot data
Stage IIIPost-training · task-specific data
LD4WAM method architecture: latent dynamics model and world dynamics action model
05 / DATA AT SCALE
270M+frames
5,000+hours of video
76.4%human data
15FPS unified format

A curated corpus spanning human egocentric videos and multi-embodiment robot demonstrations, standardized into a unified LeRobot format.

06 / VIDEO PRESENTATION
07 / EVIDENCE

Latent dynamics.
Real-world action.

LD4WAM transfers across objects, backgrounds, grippers, and dexterous hands.

ROBOTWIN AVERAGE93.4%50 tasks · clean + randomized
REAL-WORLD AVERAGE70.5%7 tasks · 2 embodiments
FULL MODEL GAIN+23.7 ptsvs. video-action baseline
REAL-WORLD CAPABILITIES

Distinct manipulation regimes, evaluated across gripper and dexterous-hand embodiments.

01

Long Horizon

Multi-step task execution

02

Dexterous Hand

Fine-grained hand control

03

Deformable

Non-rigid object manipulation

04

High Precision

Accurate insertion and placement

GENERALIZATION

Task execution remains robust when visual conditions depart from the training setting.

01

Object

Unseen object instances

02

Light

Lighting perturbation

03

Texture

Surface texture shift

04

Clutter

Distractor objects in scene

LD4WAM

Motion-aligned latent dynamics.
From video to action.

CITATION

BibTeX

@misc{shen2026ld4wamlearninglatentdynamics,
      title={LD4WAM: Learning Latent Dynamics from Human Videos for World Action Models},
      author={Zhenhao Shen and Jiaqi Liang and Jasper Lu and Feng Jiang and Yuran Wang and Chuanbo Wei and Jiayi Liu and Jianchun Yang and Qize Yu and Jiadi You and Ce Hao and Guanqi He and Chen Xie and Ruihai Wu},
      year={2026},
      eprint={2608.22403},
      archivePrefix={arXiv},
      primaryClass={cs.RO},
      url={https://arxiv.org/abs/2608.22403},
}