<- Back to news

DexTaG uses human touch to guide dexterous policy retargeting

DexTaG records full-hand tactile maps during human tool use, rewards simulated robot contact patterns during retargeting and distills a vision-proprioception controller.

Preprint · arXiv v1 · simulation and real-robot evaluation · MIT code, data and retargeter checkpoints availableSource date: Read the primary source ↗
tactile-guided reinforcement learningdexterous retargetinghuman demonstrationsopen-source robot learning
Diagram showing a tactile glove guiding simulated reinforcement learning before distillation to a tactile-free robot controller.
Original RoboSkin.ai schematic of the DexTaG training and deployment pipeline. It is explanatory artwork, not an experiment screenshot.

Researchers from the University of Massachusetts Amherst and Genesis AI released DexTaG on September 27, 2026. DexTaG records human hand motion and a 24-by-32 tactile glove map, uses those contact patterns to shape reinforcement-learning retargeting in simulation, and then distills the result into a controller that runs on a real xArm7 and WUJI hand without tactile sensing. Paper and version record.

Key takeaways

  • The dataset contains 100 marker-pen and 140 hammer demonstrations. One retargeter is trained across all trajectories for each object rather than fitting a separate policy to every demonstration.
  • In the five-seed tactile-reward ablation, trajectory completion reaches 87.5% for the marker and 76.8% for the hammer, versus 24.7% and 27.6% without tactile rewards.
  • The official MIT-licensed repository includes simulation training code and links to a 209 MB demonstration archive plus two 21 MB pretrained retargeters. Real-robot control code and student checkpoints are not described as released.

What changed

DexTaG uses human tactile readings as training guidance, not as a sensor stream for the deployed robot. A WUJI glove records five electromagnetic fingertip poses, wrist motion and a full-hand pressure map. After background correction, the authors retarget the human motion to a robot reference and align glove regions with about 700 simulated tactile points on the WUJI robot hand.

The reinforcement-learning reward has two tactile components. One uses the recorded glove map to decide which fingertips should approach the object. The other compares normalized, smoothed simulated and human contact maps. Because the maps are independently normalized, the reward matches contact pattern rather than absolute force magnitude. Tactile-guidance method.

A PPO retargeter learns marker and hammer tool use in Genesis. It observes the motion reference, simulated state, simulated tactile map and a future-reference window. Behavior cloning with DAgger then distills the teacher into a student that sees joint angles, a target object trajectory, a wrist-camera point cloud and recent proprioception/action history.

Results and the retry denominator

Across five training seeds, the complete tactile reward reaches 87.5% plus or minus 4.0 points of trajectory completion on the marker, compared with 48.6% for geometry-gated contact and 24.7% without tactile rewards. On the hammer, the reported figures are 76.8% plus or minus 27.7, 29.9% plus or minus 5.5 and 27.6% plus or minus 0.1. The large hammer variance makes the mean less conclusive than the marker result. Reward ablation.

For OakInk2, DexTaG trains one policy per object and compares it with per-trajectory ManipTrans policies. Its trajectory-weighted completion is 67.5% with one reference and 67.2% with all references. ManipTrans with the WUJI hand falls from 58.5% to 30.0%. The compute budgets are not matched: each DexTaG policy trains for 24 hours and each ManipTrans policy for eight hours on an NVIDIA L40S. Sharing policies nevertheless reduces the authors' estimated aggregate training from 1,632 to 72 policy-hours across 204 trajectories.

Real-world reporting requires special care. The primary table evaluates 30 training-set references per object with up to three attempts each and keeps the furthest stage reached. Hammer pickup, functional grasp and placement are 22/30, 17/30 and 15/30. Marker figures are 21/30, 13/30 and 10/30. Per recorded attempt, full placement is 15/83 for the hammer and 10/86 for the marker. The best-of-three percentage measures recovery with limited retries, not single-attempt reliability.

What this means for robotics

RoboSkin analysis: DexTaG offers a pragmatic answer to an embodiment mismatch. Human contact maps do not have to become robot tactile inputs; they can instead define what a plausible grasp should feel like during simulation training. That makes touch useful even when the deployed hand lacks sensors.

The limitation is feedback at runtime. After distillation, the controller cannot directly detect slip or unexpected contact, so the pipeline transfers a contact-informed behavior prior rather than delivering tactile closed-loop control. It complements, rather than replaces, sensorized dexterous manipulation.

Limitations and availability

DexTaG is an arXiv v1 preprint, and RoboSkin.ai has not reproduced it. The real-world references come from the training set, initial objects are manually placed and the headline physical result permits up to three attempts. The controller targets one xArm7-WUJI embodiment and two tools; adapting another hand requires new robot models, tactile mapping and retargeted demonstrations.

The official repository is public under MIT. Its README links a 209 MB demonstration archive and 21 MB pretrained retargeters for both tools, and documents training, evaluation and distillation. The repository notes that the tactile simulator depends on a specific Genesis fork. Readers should separately verify Google Drive file access and licenses for bundled meshes or third-party assets before redistribution.

Continue the topic

Force-control learningWrench-ACT makes force and torque the robot policy actionVisuo-tactile learningVisTacAlign turns human touch into robot training dataImpact-aware manipulationOutcome-sensitive search teaches a dexterous hand to catch with less impact