Institute for AI Industry Research, Tsinghua UniversityXiamen UniversityPeking UniversityInstitute of Computing Technology, Chinese Academy of SciencesAlibaba Group

Efficient World Action Model Inference with
Adaptive Intermediate States

Zhinan Liu1,2,† Haozhi Han3,2,† Ruge Zhang4,2 Teng Ma5 Tao Ma5
Zheng Liu5 Yifeng Chen3 Yunquan Zhang4 Ting Cao2 Yunxin Liu2 Kun Li2,‡
1Xiamen University 2Institute for AI Industry Research, Tsinghua University 3School of Computer Science, Peking University 4Institute of Computing Technology, Chinese Academy of Sciences 5Alibaba Group

† Equal contribution.‡ Corresponding author.

01 /

Introduction

Conventional WAM inference versus WAMachine: remapping, anticipatory inference with observation rebinding, and residual rescaling Enlarge figure
Figure 1 · WAMachine at a glance. Instead of restarting every inference pass from noise, WAMachine preserves and adapts useful state across replans, denoising steps, and Transformer layers. The illustrated accepted path completes k1 steps before the real observation arrives and k2 steps afterward, with k1 + k2 < K.

Original figure ↗

World Action Models (WAMs) jointly predict robot actions and future environment evolution, enabling future-aware closed-loop control and planning. However, many WAMs rely on iterative diffusion or flow inference, repeatedly recomputing highly related states across replans, denoising steps, and Transformer layers. This introduces substantial GPU computation and observation-to-action latency.

We observe that WAM inference exhibits strong state continuity: intermediate inference states often remain informative as the control loop and computation evolve. This suggests a different view of WAM inference—not as a sequence of independent solves, but as a continuously evolving stateful process in which useful computation can be preserved and adapted rather than repeatedly discarded.

Based on this insight, we introduce a training-free stateful inference framework that exploits continuity at three levels. Trajectory Remapping carries informative trajectory state across consecutive replans; Observation Rebinding advances inference during physical execution and adapts retained denoising states once the real observation arrives; and Residual Rescaling reuses temporally coherent intermediate representations across adjacent denoising steps to reduce repeated Transformer computation.

Across three representative WAM architectures on LIBERO and RoboTwin 2.0, our framework achieves 1.47–3.05× speedups in observation-to-action latency and 2.23–3.27× speedups in GPU inference time per replan, while retaining 96.69–99.54% of native task success.

02 /

Method

WAMachine framework with trajectory remapping, observation rebinding and residual rescaling, including consistency checks and refresh paths Enlarge figure
Figure 3 · WAMachine framework. Trajectory Remapping transfers useful state between replans, Observation Rebinding adapts anticipatory denoising to the real observation, and Residual Rescaling reuses checked intermediate representations across Transformer layers. Failed consistency checks refresh the corresponding state.

Original figure ↗

Start from a useful solution, not just noise.

Retain the previous replan’s denoised endpoint and a normalized denoising direction. Remap them to the new planning horizon and reconstruct an initialization at a supported solver stage, leaving fewer denoising steps to perform.

Preserve → adapt

Previous endpoint + direction → remapped initialization

Shift trajectory slots only when the model’s execution protocol requires it. Use fresh noise for uncovered slots and for the first replan.

The mechanisms work together, but refresh independently. A rejected anticipatory prefix restarts from the remapped initialization—not necessarily from noise. After observation rebinding, residual reuse still passes its own checks under the real condition. See Section 3 and Appendix B of the paper.

03 /

Results

Fast-WAM-IDM

LIBERO200 episodes · 40 tasks

GPU inference time per replanMilliseconds · lower is better

3.27×speedup

Native: 367.49 ms. WAMachine: 112.52 ms. Speedup: 3.27 times.

Observation-to-action latencyMilliseconds · lower is better

3.05×speedup

Native: 416.64 ms. WAMachine: 136.42 ms. Speedup: 3.05 times.

Table 1: Inference efficiency and subset success on a fixed 200-episode subset. Latency is reported in milliseconds, and speedups are relative to Native.
MethodGPU / replan ↓
ms
Speedup ↑
GPU
O2A latency ↓
ms
Speedup ↑
O2A
Subset SR ↑
%
Native367.491.00×416.641.00×98.0
RTI-DP269.271.36×306.641.36×63.0
RTC729.270.50×1,959.680.21×82.0
VLA-Cache636.890.58×714.840.58×96.5
BAC276.121.33×324.391.28×99.0
WAMachine112.523.27×136.423.05×97.0

Table 1 · Inference efficiency and subset success. Results use the same fixed 200-episode subset for each WAM. Latency is reported in milliseconds, and speedups are measured against Native. Lower latency and higher speedup or success are better.

Large-sample task success

Table 2 · Large-sample closed-loop task success for Native and WAMachine. Cosmos Policy and Fast-WAM-IDM use 6,000 LIBERO episodes per method; Motus uses 1,000 Clean and 1,000 Randomized RoboTwin 2.0 episodes per method.

Cosmos Policy

Native98.05%
WAMachine97.02%
Retention98.95%
Cosmos Policy success rate by suite, in percent
SuiteNativeWAMachine
Spatial97.4097.00
Object99.6099.13
Goal97.6796.47
Long97.5395.47

Fast-WAM-IDM

Native98.60%
WAMachine98.15%
Retention99.54%
Fast-WAM-IDM success rate by suite, in percent
SuiteNativeWAMachine
Spatial99.2798.80
Object99.6798.87
Goal98.4798.00
Long97.0096.93

Motus

Native86.10%
WAMachine83.25%
Retention96.69%
Motus success rate by suite, in percent
SuiteNativeWAMachine
Clean86.5084.30
Randomized85.7082.20
04 /

Analysis

Evidence for trajectory, denoising and layer state continuity: latent similarity, anticipatory prefix readiness, and residual rescaling error Enlarge figure
Figure 2 · Why state reuse works. Across consecutive replans, denoising steps, and Transformer layers, intermediate states change gradually enough to be reused after lightweight adaptation and consistency checks.

Original figure ↗

a. Replan state remains informative.

Matching early-stage latents are similar across neighboring replans. Remapped three-step inference preserves strong task success in the Cosmos Policy diagnostic.

b. Action execution offers a window.

Most anticipatory prefixes finish within the measured execution window. Rebinding tests whether their state can be continued under the realized observation.

c. Adaptation matters.

Rescaling and refreshing residual state controls output error better than repeatedly reusing the first-step residual in the Action-DiT diagnostic.

Ablation Study

Mechanism ablation · Fast-WAM-IDM / LIBERO

Table 3: Ablation of WAMachine on LIBERO with Fast-WAM-IDM. TR, OR, and RR denote Trajectory Remapping, Observation Rebinding, and Residual Rescaling.
VariantTRORRRGraphGPU / replan ↓
ms
O2A ↓
ms
Subset SR ↑
%
Native––––415.19470.9697.5
Native (CUDA Graphs)–––✓295.29350.2598.0
WAMachine✓✓✓✓114.75144.7598.0
w/o Trajectory Remapping–✓✓✓188.02177.2097.0
w/o Observation Rebinding✓–✓✓108.45165.6295.5
w/o Residual Rescaling✓✓–✓188.99205.4599.0
w/o CUDA Graphs✓✓✓–163.99185.1597.5

Table 3 · Mechanism ablation on LIBERO with Fast-WAM-IDM. TR, OR, and RR denote Trajectory Remapping, Observation Rebinding, and Residual Rescaling. Latency is reported in milliseconds; subset success uses the same fixed 200 episodes as Table 1.

Effects of individual mechanisms. Removing Trajectory Remapping or Residual Rescaling increases GPU inference time per replan from 114.75 ms to 188.02 ms or 188.99 ms, respectively. Removing Observation Rebinding increases observation-to-action latency from 144.75 ms to 165.62 ms even though GPU inference time falls to 108.45 ms. The complete WAMachine achieves the lowest observation-to-action latency among the tested variants while maintaining comparable subset success.

05 /

Citation

If you find this work useful, please cite it using the BibTeX entry below.

BibTeX
@misc{liu2026wamachine,
  title  = {Efficient World Action Model Inference with
            Adaptive Intermediate States},
  author = {Liu, Zhinan and Han, Haozhi and Zhang, Ruge and
            Ma, Teng and Ma, Tao and Liu, Zheng and
            Chen, Yifeng and Zhang, Yunquan and Cao, Ting and
            Liu, Yunxin and Li, Kun},
  year   = {2026},
  note   = {Manuscript}
}

Download .bib ↗

Paper figure