FlashDexRetarget trains one policy across many human hand motions
FlashDexRetarget amortizes physics-based retargeting across a motion collection, using object geometry, hand-contact distance and future references.

KAIST AI and Holiday Robotics researchers released FlashDexRetarget on October 1, 2026. Instead of optimizing a new controller for every human hand-object motion, the framework trains a shared reference-conditioned reinforcement-learning policy across a collection. On a 50-motion XHand benchmark, it reports 90% SPIDER success in 29 GPU-hours; the evaluated CHORD implementation reaches 46% in 2,847 GPU-hours. Paper and version record.
Key takeaways
- The main benchmark contains 25 single-object and 25 two-object motions from HOT3D, TACO and OakInk2, evaluated in Isaac Sim on RTX 3090 GPUs.
- Object point clouds, hand-to-object distance features and the next ten reference frames help one policy distinguish geometry, contact intent and motion direction.
- The 2,847-to-29 GPU-hour comparison is 98.2 times under the authors' setup. It supports the paper's “up to 100×” wording, but it is not a general speed guarantee against every retargeting system.
What changed
Physics-based retargeting asks a robot hand to reproduce the demonstrated object motion through contacts that are feasible for its own kinematics. Single-motion optimization repeats that process for every clip. FlashDexRetarget distributes reference trajectories across parallel environments and trains one off-policy controller. Successful rollouts become robot trajectories, so optimization cost is shared across motions. Method and evaluation protocol.
Each wrist-local observation includes 128 surface points for the current simulated object and next reference pose. Signed-distance features describe fingertip and wrist proximity to the object. A temporal encoder compresses hand and object states from the next ten reference frames into a 128-dimensional vector. Separate left- and right-hand actor-critic pairs receive hand-specific rewards, which matters when each hand manipulates a different object.
The training algorithm adapts FlashSAC with a replay buffer enlarged from 10 million to 50 million transitions and a critic hidden dimension increased from 256 to 1,024. This is a substantial configuration, and the compute comparison includes algorithm and architecture choices rather than isolating a single trick.
Results under the reported conditions
For XHand, FlashDexRetarget reports SPIDER/ManipTrans-object/ManipTrans success of 0.90/0.86/0.86, 10.95 mm mean object-position error and 29 GPU-hours. Do as I Do reaches 0.36 SPIDER in 66 GPU-hours; DexMachina reaches 0.48 in 447; CHORD reaches 0.46 in 2,847. The 44-point gain over CHORD is on SPIDER success. For Sharpa Wave Hand, FlashDexRetarget reports 0.72/0.70/0.70 and 33 GPU-hours, versus CHORD's 0.50/0.22/0.00 and 3,314 GPU-hours.
Metric choice matters. SPIDER averages object errors and can overstate success when one object in a two-object motion barely moves. The authors therefore use the stricter object-only ManipTrans criterion for ablations. RoboSkin analysis: preserving both numbers is more informative than repeating the 90% headline alone.
Scaling tests train on 200, 500 and 1,000 references for 600 million environment steps. The project page reports more than 600 of 1,000 motions converted, but the paper does not present a universal conversion fraction across arbitrary motion distributions. Physical replay shows wiping a board, pouring into a pan and closing a lid. No trial count, failure rate or closed-loop policy adaptation is reported for those hardware demonstrations, so they establish executability examples rather than a quantitative sim-to-real benchmark.
RoboSkin analysis
FlashDexRetarget targets a practical bottleneck between human-motion collections and dexterous robot hands: a demonstration is not usable robot data until contact and object motion are feasible for a specific embodiment. Amortizing that conversion can matter more than slightly improving one hand's single-clip fit.
The future-reference encoder is also an important contact insight. One pose cannot distinguish whether a hand is about to wipe, pour or close. By exposing the upcoming motion and object geometry, the controller can prepare for contact rather than chase it frame by frame. That complements tactile data collection even though the method itself uses simulated geometry and state rather than physical tactile sensor input.
Limitations and availability
FlashDexRetarget is an arXiv v1 preprint, and RoboSkin.ai has not run the code or reproduced the metrics. Main results are simulation-based, use three source datasets and two robot hands, and compare methods under the authors' selected preprocessing and success definitions. Real-world evidence is qualitative replay of three tasks.
The official project publishes videos, method details and tables but labels code “coming soon.” No repository, trained policy, converted trajectory archive or implementation license was verified on October 2, 2026. The paper is CC BY 4.0; that license covers the article, not unreleased software or third-party source datasets.


