SkeleWAM compresses manipulation into a sparse 3D skeleton
SkeleWAM replaces visual-future reconstruction with robot joints, object centers and interaction points, using future geometry as training-only supervision.

Peking University researchers released SkeleWAM on October 1, 2026, asking whether a world action model needs to predict pixels at all. The model converts RGB-D observations and robot proprioception into a sparse 3D skeleton of robot joints, object centers and interaction points. It jointly learns actions and future skeletons during training, then removes the future-prediction branch at inference. Paper and version record.
Key takeaways
- The RGB-D model has about 57.1 million parameters and reports 85.9% on 10,030 LIBERO-Plus perturbation variants, 3.7 points above the compared Cosmos-Policy result.
- Removing future-skeleton supervision reduces success from 85.9% to 80.1%, even though both variants use the same action-only inference path.
- On five ARX R5 tasks with 20 trials each, SkeleWAM records 89/100 successes; the margin over Cosmos-Policy is two percentage points.
A geometry-first world state
Many world action models predict a future video or a latent produced by a visual encoder. Those targets can preserve appearance that is unrelated to control while leaving contact geometry implicit. SkeleWAM instead builds object centers and manipulation-relevant interaction points from RGB-D, adds robot keypoints from forward kinematics and expresses them in a shared robot-centric frame.
A shared world expert processes the current skeleton and language. Separate flow-matching heads predict a 32-step action chunk and eight future skeleton states. The future head is training-only. At inference, Medoid Action Consensus samples three trajectories, compares their first ten motion steps and executes the representative candidate rather than averaging possibly incompatible actions. The default controller executes 16 actions before rebuilding the skeleton. Architecture and ablations.
Results and comparison boundaries
The main simulation benchmark covers seven perturbation categories. SkeleWAM reaches 93.4% under camera changes, 71.9% for robot initial-state changes, 89.1% for language, 94.7% for lighting, 96.0% for background, 93.9% for noise and 66.6% for layout, averaging 85.9%. Cosmos-Policy reports 82.2% overall and 75.8% for camera perturbations. However, π0.5 reaches 84.1% on layout changes versus SkeleWAM’s 66.6%; sparse geometry does not eliminate spatial-generalization failures.
The privileged sim-state skeleton reaches 87.7%, only 1.8 points above RGB-D overall, but still only 69.0% on layout. That indicates perception error is not the sole cause of the layout weakness. The policy itself must generalize to different spatial arrangements.
Physical evaluation uses an ARX R5 with external and wrist-mounted RealSense cameras. Across Open Drawer, Close Drawer, Stack Blocks, Stack Bowls and Put Block in Drawer, SkeleWAM reports 80%, 90%, 90%, 95% and 90%, respectively—89 successes across 100 trials. Cosmos-Policy totals 87/100, Fast-WAM 83/100 and π0.5 76/100 under the authors’ setup. This is a useful hardware check, but the two-success margin over Cosmos-Policy is small and no confidence interval is supplied.
RoboSkin analysis
SkeleWAM’s editorially important result is the ablation, not only the leaderboard. Centers alone score 82.3%; interaction points alone score 77.7%; combining them scores 85.9%. Static location and task-relevant contact geometry carry complementary information. That is a practical design clue for robot world models and contact-rich manipulation: a compact state should preserve both where an object is and where the policy can act on it.
The system is not tactile. Interaction points are inferred from RGB-D rather than measured through touch. A visuo-tactile world model could use the skeleton as a geometric backbone while adding force, slip or contact-state evidence that cameras cannot recover under occlusion.
Limitations and availability
The manuscript is an arXiv v1 preprint, and RoboSkin.ai has not reproduced it. Training and evaluation use one RTX 4090 in the reported implementation, but training time and end-to-end control latency are not given. Hardware evidence covers one arm, five tabletop tasks and 20 trials per task. The benchmark’s strongest weakness is layout change, and the real-world comparison is too small to establish broad superiority.
The official project provides figures, result tables and recorded demonstrations. No implementation repository, checkpoint, training dataset, perception model package or software license was verified. The paper is posted under CC BY-NC-ND 4.0, which covers the manuscript, not an absent software release.


