<- Back to news

TacDyn-WAM predicts contact dynamics without generating future touch pixels

TacDyn-WAM separates visual futures from implicit tactile dynamics, replacing iterative tactile-pixel generation with a contact-aware latent target.

Preprint · arXiv v1 · eight-task simulation benchmark · 100 physical trials per evaluated policy · project videos public · code coming soonSource date: Read the primary source ↗
tactile world action modelimplicit contact dynamicsvisuo-tactile manipulationUniVTAC
Diagram showing camera and tactile clips entering separate future-prediction experts before guiding a robot action policy.
Original RoboSkin.ai schematic of TacDyn-WAM's separate visual and tactile prediction paths. It is explanatory artwork, not an experimental figure.

Researchers led by Tsinghua University's Institute for AI Industry Research released TacDyn-WAM on September 30, 2026. The world action model predicts how contact representations will evolve instead of reconstructing future tactile images through iterative denoising. It reports 81.5% average success on eight UniVTAC tasks using the benchmark demonstrations, then 71.0% across five physical tasks without its extra pretraining stage and 85.0% with that stage. Paper and version record.

Key takeaways

  • TacRep learns a tactile target space from four-frame clips; an Implicit Tactile Dynamics Expert predicts future representations and their changes at several horizons in one forward pass.
  • A separate read-only tactile memory describes current contact. The action expert can therefore use both present touch and a forecast of contact evolution.
  • The real-robot study uses five tasks, 60 demonstrations per task and 20 trials per policy per task. The reported 85.0% is 85 successes per 100 trials when the published per-task percentages are summed.

What changed

Most tactile world models inherit a video-generation objective. That preserves spatial detail, but the target can be brittle: a small shift in contact location may alter many tactile pixels even when the physical trend—deepening, sliding or rotating—remains predictable. TacDyn-WAM assigns vision and touch different target spaces. A visual expert predicts future visual latents, while the tactile expert predicts TacRep features and feature changes. Joint attention lets the two experts exchange information without forcing touch into a camera-oriented reconstruction space. Architecture and training stages.

The tactile loss covers 49 patches per sensor and upweights patches whose representation changes more. A compact understanding path also compresses a frozen AnyTouch2 encoding into ten tokens. The model is trained in stages: tactile representation learning, tactile-world grounding, tactile–action alignment and joint training.

Results under the reported conditions

On UniVTAC, each of eight tasks is evaluated with 100 rollouts. TacDyn-WAM averages 81.5%. The paper groups the 83.1% N0-VTLA and 84.5% N0-TWAM results separately because those systems use large-scale visuo-tactile trajectory pretraining. TacDyn-WAM uses the provided per-task demonstrations for this comparison; it is therefore close in outcome, but not evidence that the systems have equal training cost or generality.

Ablations give the clearest evidence for the target design. Replacing TacRep with the base model's reconstruction-oriented Cosmos VAE reduces the average to 62.0%; a static DINOv2 target reaches 74.0%. Removing the tactile world model gives 67.9%, while removing current-state tactile memory gives 71.3%. These are absolute percentage-point comparisons within the authors' common protocol.

For hardware tests, a Franka Research 3 uses Xense sensors on both gripper fingers plus wrist and third-person cameras. Across Stack Cups, Remove Plug, Insert Plug, Unscrew Cup Lid and Wipe Whiteboard, the non-pretrained model records 71/100 successes. Pretraining on a 6,000-trajectory OmniViTac subset, with the 300 task demonstrations also used in early stages, raises the total to 85/100. On one A100, the paper reports 484 ms for a 50-action chunk versus 1,338 ms for an optimized 40-action LingBot-VA chunk; that is a system-specific latency comparison, not a universal real-time guarantee.

RoboSkin analysis

The useful engineering idea is not simply “latent is faster.” Contact images can be visually different while encoding the same corrective direction. A dynamics-aware target can focus capacity on what the controller needs next. That makes TacDyn-WAM relevant to visuo-tactile world models and to teams designing a tactile manipulation loop with distinct perception and action timescales.

The benchmark comparison also needs care. The 81.5% simulation average trails the full N0 models, and the physical gains come from one vision-based tactile sensor family. The paper explicitly leaves force sensors and taxel arrays for future work. Cross-sensor robustness is therefore unproven.

Limitations and availability

TacDyn-WAM is an arXiv v1 preprint and RoboSkin.ai has not reproduced the results. Hardware evidence covers one parallel gripper, five tasks and randomized but laboratory-controlled starts. The official project publishes tables and a downloadable demo video, but labels code “coming soon”; no public training code, weights, real-robot dataset or implementation license was verified. The paper's arXiv license does not supply those missing rights.

Continue the topic

Contact-aware world modelsInternW0 links asynchronous world prediction with contact-aware pipettingTactile robot controlAgile-WAM pairs fast touch prediction with robot actionsTactile AITouchWorld separates tactile prediction from fast contact correction in robot manipulation