World action models · driving and manipulation

CtrlWAM

Controllable World Action Models with
Aligned Intent and Foresight

Chensheng Peng1,2, Wenhao Ding2, Ran Tian1, Zewei Zhou2,3, Jef Packer2, Maximilian Igl2, Peter Karkus2, Yan Wang2, Masayoshi Tomizuka1, Boris Ivanovic2, Marco Pavone2, Yuxiao Chen2

1University of California, Berkeley  ·  2NVIDIA  ·  3University of California, Los Angeles

Driving on NVIDIA PhysicalAI-AV + NuRec · Manipulation on RoboTwin 2.0
1.023m
SFT baseline 1.023
0.538m
SFT baseline 0.538
10.07%
ego-only 10.07 %
58.70
SFT 58.70
scroll

Overview

Intent and foresight
physically aligned
One world action model

World action models jointly predict an agent's future actions (intent) and the video of what it will see (foresight). CtrlWAM trains them so that noised actions and noised video stay physically consistent, and lets a single model predict or follow commands for the ego agent and the agents around it.

Read the abstract

World action models (WAMs) jointly predict actions (intent) and visual future (foresight). Standard training adds noise to recorded actions and video simultaneously, but such training paradigms introduce a mismatch: perturbed actions imply counterfactual future visual, while the noised video remains tied to the GT recording. In low-noise regime, the scene geometry and even the dynamic behavior remain clearly visible from the noisy future frames despite the added noise.

We present CtrlWAM, which executes perturbed actions in a simulator and pairs them with their noised visual consequences for joint WAM learning. To accommodate the different denoising requirements of video and actions, we introduce warped video–action noise schedules that aim to keep visual layout responsive as action predictions evolve. We further extend the action interface from ego-only control to a variable number of agent streams, allowing a unified model to represent predicted or commanded futures for multiple agents.

Driving experiments show more accurate action forecasts, closer agreement between generated video and actions, and better following of supplied commands; robotics experiments show stronger motion fidelity and controllability. Matched controls support the benefit of off-path renders for command following and manipulation fidelity. Together, these findings contribute to a more controllable world action model.

CtrlWAM at a glance. Left: perturbed actions and their rendered consequences are corrected toward the recorded pair, with independent noise schedules for actions and video. Middle: one model covers joint, video and action prediction, and controls agents beyond the ego. Right: the same recipe applies to driving and robotic manipulation.
Joint prediction

Actions + video, together

Both future streams are denoised jointly; the imagined video should reflect the predicted intent.

Forward dynamics

Action-conditioned video

Future actions are supplied clean as commands; only video is generated. This is where controllability is measured: does the video follow a rewritten command?

Inverse dynamics

Video-conditioned actions

Clean future video, noised actions. A separately trained inverse dynamics model (IDM) also reads trajectories out of generated video for evaluation.

Motivation

Standard noising pairs a perturbed action with video of the recorded future

Three states, replaying. Perturbing the action changes the implied future — but Gaussian noise on the recorded frames never moves the car.

(a) CleanRecorded pair · safe pass
t₁
✓ video: safe pass video ≈ recorded future, action perturbed ✕ action: collision with the neighbour
The model recovers the clean pair, but gets no supervision on the physical correspondence between its intermediate action and video states.
(a) Clean

The recorded pair agrees — at the clean endpoint only

A recorded action–video pair specifies how the two modalities agree at σ = 0. It says nothing about how they should deviate together as noise is added.

(b) Noised

Perturbed action, unchanged video

Methods such as DreamZero corrupt actions and video with independent Gaussian noise. The noised action implies a different ego path and a different visual future — but the corrupted video remains tied to the recording. In the low-noise regime the scene layout and even the dynamics stay clearly visible.

(c) Denoised

Safe video, unsafe action

Useful visual structure emerges early in denoising (DIDO: step-1 video features already give 97.7 % vs 98.4 % after four steps on LIBERO). Without alignment, an intermediate action estimate can veer toward an obstacle while the imagined video still depicts a safe pass.

Video commits to a layout early; the action keeps changing

In both driving and manipulation the visual future is recognisable while σ is still high — the explicitly predicted trajectory is still moving. Under a shared schedule the video can therefore commit before the action settles.

pure noise
clean
joint generation traces, 35 / 30 UniPC steps · replays automatically — drag the bar to hold a step

Standard training noises video and action with one shared level (red diagonal): both are assumed equally uncertain at every step.

The animation above shows otherwise. The action the partially denoised video implies (blue) converges within the first few steps, while the video itself is still far from clean.

The horizontal gap between the curves is the denoising mismatch: the action token is trained against a video that already fixed the ego path, yet nothing in the pair tells the model the two must agree along the way.

Method

Three components, one training recipe

Explore how CtrlWAM constructs physically corresponding inputs, paces video against actions, and represents surrounding agents.

Pair every perturbed action with its rendered consequence

Standard noising vs aligned noising. Standard noising perturbs the action but leaves the recorded video in place: the noised action and the noised video disagree (✕). Aligned noising executes the perturbed action in a simulator, renders the views it produces, and noises those renders, so action and video agree (✓); both are still trained toward the recorded pair in the middle. Frames show the driving and manipulation domains together; press ↺ to replay.

Left: standard noising keeps the recorded future in the video. Right: the simulator executes the noised command, its render is noised, and both branches are still trained toward the recorded pair.

  1. Sample the level. Draw t\in[0,1), set \sigma_a = 1-t, and form the perturbed command \mathbf{a}^{\mathrm{cmd}}_t = t\,\mathbf{a}^\star + (1-t)\,\boldsymbol{\epsilon}_a.
  2. Execute and render. A simulator runs the command — AlpaSim on NuRec 3DGS scenes for driving, RoboTwin 2.0 for manipulation. We store the executed trajectory \mathbf{a}_t and the render latents \mathbf{r}.
  3. Noise the render. \mathbf{v}_{\sigma_v} = \sigma_v\,\boldsymbol{\epsilon}_v + (1-\sigma_v)\,\mathbf{r}. The input now pairs a perturbed action with a corrupted view of its consequences: physical correspondence.
  4. Supervise toward the recording. Targets are the recorded actions and video at the same timestamps, so the model learns to correct both the behavioural and the visual discrepancy.

Hold the video at higher noise while the actions settle

\sigma_v=\mathrm{warp}(\sigma_a;\varsigma)=\dfrac{\varsigma\,\sigma_a}{1+(\varsigma-1)\,\sigma_a}
  • \varsigma=1 recovers the shared schedule \sigma_v=\sigma_a.
  • For \varsigma>1, at a fixed action noise level the video input carries more Gaussian noise and less rendered content; the two levels stay deterministically coupled.
  • Each predicted modality gets its own timestep embedding; actions supplied as conditioning stay fixed.
  • Intended effect: leave room for changes in scene layout while the action is uncertain, then let the video settle into detail refinement as the action converges.

Predict or command every agent in the scene

  • One ego block plus K agent blocks, each with history and future tokens (14 + 64 rows at 0.1 s).
  • No commands → the model jointly predicts ego, agents and video; predicted agent futures act as intent estimates that inform the ego action.
  • Commands → the agent's future tokens are supplied clean; the model predicts the video (and any remaining agents).
  • K=0 recovers the ego-only model exactly.
Encoding

Agent histories store ego-frame positions at t_0 (scaled by 20 m); futures are unicycle controls in each actor's local frame. Actors are selected by visibility in the front camera (≤ 60 m, ≥ 12 px, not occluded), up to K_{\max}=8; a shared agent projection and per-actor spatial rotary positions place them on the latent grid.

Model

One diffusion transformer, three token streams

A shared transformer processes video, ego-action and optional agent-action tokens, conditioned on observed history. Future tokens are denoised as prediction targets or supplied clean as conditions — the same network serves joint prediction, action-conditioned video and video-conditioned actions.

Tokens flow through the model: noised future tokens from each modality enter the transformer, attend jointly with the clean history condition, and are decoded back into video, ego action and agent trajectories.

Video

Future frames → causal VAE (4× time, 16× space) → latent tokens. Noised in forward dynamics and joint mode, clean for inverse dynamics.

Ego action

Unicycle controls (acceleration, curvature) or 14-D joint positions; 64 future rows. Noised for prediction, clean when supplied as a command.

Agents (optional)

0…K agent blocks through a shared agent encoder/decoder; variable length. Backbone: Cosmos 3-Nano (16B), language pathway frozen.

Results · driving

Better command following, consistency and forecasting

1,156 held-out clips from PhysicalAI-AV. In Joint mode the model predicts video and actions together; in forward dynamics the future actions are supplied as commands. CtrlWAM-S is the ego-only model, CtrlWAM-M the multi-agent model, trained with the same recipe, data and budget.

Joint modeForward dynamics
ModelPSNR (dB)Self ADE (m)Rollout ADE (m)PSNR (dB)Cmd ADE (m)Cmd FDE (m)
Cosmos 3-Nano (zero-shot)19.586.2128.26720.051.6004.040
SFT baseline20.380.5382.92021.231.0232.803
CtrlWAM-S (ego-only)20.340.5212.84821.231.0012.722
CtrlWAM-M (multi-agent)20.420.5132.80421.280.9812.671

Self ADE compares the IDM read-out of predicted video with the model's own action trajectory (video–action consistency). Rollout ADE compares the generated actions with the recording. Cmd ADE/FDE compare the IDM read-out with the supplied command. Bold = best, underline = second best.

Results · counterfactual ego control

Change the command, keep the scene — the video follows

Scene history, text prompt and sampling noise are held fixed; only the ego command changes. Stopping halts forward viewpoint progression, acceleration advances the ego past the parked cars.

Explore the clip
t = 0.0 s
Command deviation from the recordingrecorded
IDM read-out vs command (Cmd ADE)0.91 m
PSNR vs the real video19.9 dB

Top-down inset: the commanded ego path in the t₀ frame (10 m rings), magenta when rewritten. Generated by the multi-agent CtrlWAM model in forward dynamics; one clip at 10.7 m/s, 5 Hz frames over 6.4 s.

Warped schedules follow modified commands more closely

  • 100 held-out clips; each level averages four rewritten commands: brake to stop, accelerate, nudge left, nudge right.
  • Severity = how far the command deviates from the recording: 4 / 7 / 11 m (low / mid / high). Low / mid / high use decelerations of 1 / 2 / 3 m/s², accelerations of 0.75 / 1.5 / 3 m/s² and lateral offsets of 2 / 3.5 / 5 m.
  • Warped schedules (ς = 2, 3, 5) beat the shared schedule at every level, and the gap widens as perturbations grow.
  • Training on GT pairs (real frames, no off-path renders) degrades fastest: the rendered consequences are what teach command following.

Results · multi-agent control

A dedicated channel for commanding surrounding vehicles

Scene history, the ego command and sampling noise are held fixed while a selected agent's future trajectory is rewritten. Gemini 3.8 Flash rates agreement between the generated video and the supplied agent command (0–1, reported in %). Only CtrlWAM-M can receive the command; the other models are scored on agreement without conditioning.

Counterfactual agent control

Counterfactual control of the lead vehicle. Top: factual agent trajectory. Bottom: the same scene with the agent commanded to stop — the generated video shows it halting in front of the ego. Right: unchanged ego path (grey), factual (green) and rewritten (magenta) agent command. Mean compliance rises from 10.07 % (ego-only) to 18.94 %.

Results · robotic manipulation

Highest score in all six WorldArena categories

RoboTwin 2.0 bimanual manipulation, 1,000 held-out episodes, WorldArena Track-1. Each arm starts from a common 2,000-iteration checkpoint and receives 2,000 more iterations using aligned renders (CtrlWAM), real frames in place of renders (GT-Src) or standard SFT. EWMScore is 100 × the mean of the 15 individual metrics.

CategoryCosmos 3
zero-shot
Cosmos 3
SFT
GT-SrcCtrlWAMOracle
Visual quality0.4470.5600.5630.5760.598
Motion quality0.2470.3160.3060.3270.348
Content consistency0.4130.6500.6200.6650.662
Physics adherence0.2580.4560.4800.5400.782
3D accuracy0.8120.8840.8990.9210.956
Controllability0.6350.7740.7880.8190.869
EWMScore44.8858.7058.6561.7566.92

Bold = best, underline = second best; Oracle scores the recorded videos and is excluded from ranking. Individual metrics do not uniformly favour aligned training (continued SFT has slightly higher image and aesthetic quality); the aggregate gain comes from JEPA similarity, trajectory accuracy, interaction quality and instruction following.

Explore episodes

Conclusion

Physically aligned denoising for controllable world action models

CtrlWAM makes the relationship between an action deviation and the observation it produces part of the denoising supervision. Warped schedules keep visual layout responsive while actions evolve, and multi-agent streams allow surrounding futures to be predicted or supplied as commands.

Limitations. Aligned noising needs a simulator that can execute perturbed actions; the multi-agent extension is shown for driving only; offline metrics do not establish closed-loop safety.

Citation

@misc{ctrlwam2026,
  title = {CtrlWAM: Controllable World Action Models with Aligned Intent and Foresight},
  year  = {2026},
  note  = {Preprint}
}