Native Action-Prior Learning from Videos
for World Action Models

NAVA-WAM pretrains the action policy directly from observation-only videos — no intermediate visual representation, no separate latent-action model.

1Meta AI 2University of Copenhagen 3Imperial College London 4Physical Intelligence
†Work done at Meta*Project lead and corresponding author

Real-Robot Deployment

The DROID-post-trained policy runs on a physical Franka FR3 without any task-specific fine-tuning. Each clip shows the third-person view, with the wrist camera inset.

4/5 T1 — move the cube to the left side of the bowl
5/5 T2 — put the banana in the box
5/5 T3 — put the blue cube on the red cube
Modelms / callT1T2T3Avg (%)
π0.5195.70/55/53/553.3
DreamZero3523.90/55/55/566.7
NAVA-WAM (ours)379.24/55/55/593.3

Five trials per task, 15 trials total. ms / call is server-side latency per policy call on a single H100; normalizing by the number of actions executed per replanning step gives 15.8 ms per executed action for NAVA-WAM, versus 24.5 ms for π0.5 and 146.8 ms for DreamZero.

T1: grounding the spatial relation

The largest gap is on T1, which requires placing the cube to the left of the bowl. NAVA-WAM succeeds in 4/5 trials while both baselines fail all five. In all ten baseline trials the cube is localized and grasped correctly — the failure is in grounding the instructed spatial relation, not in perception or manipulation.

success NAVA-WAM (ours) — places the cube to the left of the bowl
failure DreamZero — ends with the cube in the bowl or still held above it
failure π0.5 — places the cube to the right of the bowl, or inside it

Abstract

World action models integrate future visual dynamics with robot action prediction, but their scalability remains limited by the need for action-annotated robot trajectories. Observation-only videos contain rich evidence about interaction dynamics, but existing approaches typically use them either to pretrain visual representations that must later be adapted for control, or to infer latent actions that are subsequently grounded to robot commands.

We present NAVA-WAM, which introduces native action-prior learning by directly pre-training the action policy from observation-only videos, avoiding indirect representation-to-control transfer or a separate latent-action model. Our training consists of two stages. First, we pretrain on observation-only videos, where future-video flow-matching supervision over visual transitions is propagated through transition-structured joint attention to optimize the Action-DiT and learn action-relevant priors. Second, we use action-labeled demonstrations to post-train the Action-DiT for robot control through joint video–action flow matching, while asymmetric attention decouples the visual branch from iterative action denoising and enables efficient action-only inference.

Extensive experiments show that NAVA-WAM consistently outperforms prior approaches under both in-distribution and out-of-distribution evaluation, while demonstrating strong action-label efficiency and effective real-robot generalization. These results establish native action-prior learning as an effective approach to directly pretrain action policies from observation-only videos, providing a scalable path beyond action-labeled robot data.

Method

Overview of NAVA-WAM: action-free pre-training and downstream post-training.
Overview of NAVA-WAM. During action-free pre-training, the Action Expert is optimized solely through future-video supervision, learning action-relevant priors directly from observation-only videos. During post-training, both experts are jointly optimized on action-labeled robot demonstrations using video and action flow-matching objectives, grounding the pretrained Action Expert to continuous robot control. At inference, NAVA-WAM denoises only robot actions without explicitly generating future videos, enabling efficient closed-loop control.

Stage 1 — Native action-prior pre-training

For each visual transition, we build a segment holding the clean preceding state, a noised future state, and a group of Action-DiT tokens. A transition-structured joint attention mask lets the Action-DiT read both visual states (an implicit inverse-dynamics path) while the noised future tokens attend back to the Action-DiT representations (an implicit forward-dynamics path). Because the Video-DiT is frozen, minimizing the future-video flow-matching loss updates the Action-DiT only through this joint-attention pathway. The Action-DiT sees the future only in its noised form, which rules out the shortcut of copying the clean prediction target.

Stage 2 — Post-training and action-only inference

Action-labeled demonstrations adapt the pretrained Action-DiT to continuous control via joint video–action flow matching. Here we make the cross-stream attention asymmetric: the Video-DiT no longer attends to the action stream, so its keys and values can be computed once and cached. After a single joint forward pass, iterative denoising runs through the Action-DiT alone, giving closed-loop control without ever rendering a future video.

Attention masks used during pre-training, post-training, and inference.
Transition-structured attention masks used during pre-training and post-training.

Simulation Results

NAVA-WAM maintains strong in-domain control while substantially improving out-of-domain generalization.

LIBERO / LIBERO-Plus
MethodLIBEROLIBERO-Plus
Direct action policies
π094.153.6
π0-FAST85.561.6
StarVLA-α96.577.0
π0.596.977.4
ABot-M098.680.5
World action models
JEPA-VLA96.425.6
Fast-WAM97.651.5
Image-WAM98.483.1
Being-H0.799.282.1
NAVA-WAM (ours)99.083.5
RoboTwin 2.0
MethodCleanRandom
Direct action policies
DP28.00.6
RDT34.513.7
π046.416.3
UP-VLA52.915.2
World action models
Fast-WAM71.96.3
BagelVLA75.320.5
HALO80.526.4
Image-WAM85.037.6
MV-WAM84.055.7
NAVA-WAM (ours)88.573.6

Average success rate (%). LIBERO and RoboTwin Clean measure in-domain performance; LIBERO-Plus and RoboTwin Random evaluate out-of-domain generalization. On RoboTwin 2.0, NAVA-WAM improves over the strongest baseline by 4.5 points on Clean and 17.9 points on Random, with a much smaller Clean-to-Random drop. All RoboTwin training uses Clean-domain demonstrations only.

What Does the Action Prior Actually Learn?

Motion transfer across domains

Qualitative visualization of motion transfer across domains.
Motion transfer across domains. Each example transfers the transition encoded by a source pair (Source: current → Source: next) to a different target observation. The source transition specifies the underlying motion cue (orange arrows), which is applied to Target: current to obtain Target: transfer. Overlay shows the transferred state relative to the current target observation, and Target: real shows the corresponding real transition (blue arrows). Across simulation, real-world, and sim-to-real examples, the transferred states follow the source motion despite substantial changes in appearance and scene configuration.

Attention localization

Action-DiT attention concentrates on the gripper and manipulated objects.
Attention of learned action representations. We visualize attention from action representations to visual tokens on held-out data, overlaid on the current and next input observations. Our pretrained Action-DiT attends to interaction-relevant regions and aligns with observed motion, while the CoMo latent-action baseline shows more diffuse attention over static regions.

Qualitative Rollouts on RoboTwin 2.0

All episodes below come from the Random (out-of-domain) split, evaluated with a policy trained only on Clean-domain demonstrations.

Each clip shows the head camera (left) and the two wrist cameras (right).

Successes

Handover Mic — grasp the microphone with one arm and hand it over to the other arm
Lift Pot — use both arms to lift the pot
Place Dual Shoes — put both shoes into the shoebox, with the shoe tip pointing left
Place Bread Skillet — grab the bread with one arm and put it into the skillet
Open Laptop — use one arm to open the laptop
Blocks Ranking RGB — place the red, green and blue blocks in a row, in that order from left to right
Stack Blocks Three — stack the blue block on the green block and the green block on the red block

Failures

failureMove Can Pot — pick up the can with one arm and move it beside the pot
failureHanging Mug — pick and re-orient the mug with the left arm, then hang it on the rack with the right arm

The two representative failures occur during fine-grained grasping and contact-rich manipulation. In Move Can Pot, the reaching motion knocks the can over instead of establishing a stable grasp; the arm then executes the transport motion with an empty gripper without re-attempting the grasp, and the episode eventually reaches its step limit. In Hanging Mug, the policy completes the initial pick sub-goal, lifting the mug with the left arm and placing it beside the rack, but the right arm fails to continue the task by lifting the mug onto the rack, and the episode eventually reaches its step limit.

Together, these cases suggest that failures can arise during fine-grained, contact-rich interactions even when the target object and overall task progression are correctly identified. In both examples, the policy fails to recover from a local execution error, causing it to propagate into task failure.

BibTeX

@article{an2026navawam,
  title   = {Native Action-Prior Learning from Videos for World Action Models},
  author  = {An, Zhaochong and Zhang, Fei and Jia, Menglin and Frost, Duncan and
             Zhou, Zijian and Wang, Yikai and Wang, Xudong and Patel, Aditya and
             Zeng, Belinda and Xiang, Tao and Belongie, Serge and Bar, Amir and He, Sen},
  journal = {arXiv preprint arXiv:2610.03391},
  year    = {2026}
}