WAM-Cache: Staleness-Bounded KV Reuse for Efficient World Action Models

Kai Ding1Yang He2Ruijie Quan1Yi Yang1

1Zhejiang University2IAIC, Agency for Science, Technology and Research, Singapore

dingkai@zju.edu.cnhe_yang@a-star.edu.sgquanruijie@zju.edu.cnyangyics@zju.edu.cn

Abstract

World Action Models (WAMs) enable generalist robot manipulation by conditioning an action expert on representations from a pretrained video Diffusion Transformer (DiT). In closed-loop control, the video DiT runs at every chunk to encode the current observation into layerwise key–value (KV) pairs that the action expert queries. This prefill dominates the per-chunk computational cost, yet existing training-free accelerations leave it fully dense. We present WAM-Cache, a training-free framework that retains layerwise key–value representations across chunks and recomputes only a sparse refresh set of tokens. Crucially, we find that the intuitive heuristic of refreshing visually drifted tokens plateaus far below the dense baseline, even with an oracle predicting ground-truth KV drift. Downstream action accuracy is instead governed by where the action expert attends, not by what moved. WAM-Cache therefore selects the refresh set by uniting the action expert’s cross-attention with visual latent surprise, complemented by a strict age bound that suppresses compounding error. On Fast-WAM, WAM-Cache cuts video DiT prefill FLOPs by 32–42% across RoboTwin 2.0, LIBERO, and real-world experiments, while staying within 0.7–1.8 percentage points of the dense policy in simulation and 2.5 points on a real robot.

The WAM-Cache pipeline. Panel (a) builds the refresh set in three steps: latent surprise, action-conditioned selection and age-bounded refresh. Panel (b) runs a sparse video DiT prefill over the refresh set and updates the layerwise key-value cache that the action expert reads.

Scroll sideways to see the whole figure.

The WAM-Cache pipeline. (a) Three steps build the refresh set Rt: 01 latent surprise, 02 action-conditioned selection and 03 age-bounded refresh. (b) The video DiT runs a sparse prefill over Rt and updates the layerwise keys and values in place. The action expert then denoises the next action chunk over the mixed cache.

The three steps on a real chunk

Here are steps 01 to 03 of (a) on chunk t of one RoboTwin 2.0 episode. Between chunk t−1 and chunk t, only the arm moves.

Model input at chunk t minus 1. The left arm enters from the top left, above and to the left of the blue cup. Model input at chunk t. The gripper is around the blue cup.

Surprise scorelow to high

Not refreshed, age +1

RecomputedReused from cache

RoboTwin 2.0 · place_empty_cup

Efficiency

WAM-Cache comes in two settings. The default cuts the prefill FLOPs by 41.6% on RoboTwin 2.0 and 32.3% on LIBERO, within 1.8 and 0.65 points of the dense policy. The larger budget still saves 28.2% and 26.1%. It comes within 0.1 points of dense on RoboTwin 2.0 and matches it on LIBERO.

RoboTwin 2.0
1.053 T · SR 90.16
0.615 T · SR 88.36 · −41.6% FLOPs
0.756 T · SR 90.08 · −28.2% FLOPs
LIBERO
0.859 T · SR 96.50
0.582 T · SR 95.85 · −32.3% FLOPs
0.635 T · SR 96.50 · −26.1% FLOPs
00.40.81.2Prefill TFLOPs / chunk

SR is the success rate in %. RoboTwin 2.0 uses clean scenes, and LIBERO is the average over its four suites. The default uses (ba, bs) = (0.4, 0.1) on RoboTwin 2.0 and (0.5, 0.1) on LIBERO. The larger budget uses (0.6, 0.1) on both.

BibTeX

@misc{ding2026wamcache,
  title  = {{WAM-Cache}: Staleness-Bounded {KV} Reuse for
            Efficient World Action Models},
  author = {Ding, Kai and He, Yang and Quan, Ruijie and Yang, Yi},
  year   = {2026}
}