Outcome-sensitive search teaches a dexterous hand to catch with less impact
A simulation-only study searches the short pre-contact motion window where small action changes alter impact. Its imitation policy beats the privileged teacher, but physical contact transfer remains untested.

Researchers spanning Taiyuan University of Technology, Hong Kong Polytechnic University, Southern University of Science and Technology, Great Bay University, KTH and Chery's Kaiyang Laboratory released Outcome-Sensitive Motion Search on September 24, 2026. The simulation method edits a short pre-contact motion window where small action changes strongly affect both interception and impact, then uses the successful rollouts to train a deployable imitation policy. Paper and version record.
Key takeaways
- The final imitation policy reaches 90.3% task success, 5.1 percentage points above the privileged reinforcement-learning teacher and 6.5 points above success-only imitation.
- Peak simulated impact is 37.3 N for the full method, 8.6% below the teacher's 40.8 N, under the paper's defined catch and impact thresholds.
- Every result is in MuJoCo. The paper explicitly says robustness to real perception and contact mismatch remains unverified without a physical arm-hand experiment.
What changed
A teacher that catches an object is not automatically a good source of demonstrations. It may fail on some launch conditions, and two motions that both complete a catch can create very different relative speed, impact force and follow-through.
The authors define an outcome-sensitive window around the motion immediately before contact. They learn a task-conditioned manifold of successful eight-step arm-action snippets, project teacher motions into it, and search along local geodesic directions. Search can refine a successful catch to reduce impact or repair a teacher failure. Candidate snippets are inserted back into a full rollout and kept only if the complete catch succeeds under a modeled imitation-student action error.
The simulation uses a 7-DoF xArm7 and a 17-DoF ORCA Hand at 50 Hz. Objects vary across ten shape-size combinations, 40 to 60 g mass and roughly 2.8 to 5.0 m/s launch speed. Three source datasets contain 20,000 successful, 20,000 successful and 20,000 mixed teacher rollouts respectively. The constructed training set contains 99,328 successful demonstrations.
Results under the reported conditions
The study separates catch completion from impact-aware task success. A catching-only teacher completes 92.5% of catches but satisfies all impact criteria in just 0.4%, with 160.1 N mean peak force. Adding impact objectives gives the privileged teacher 90.7% completion, 85.2% task success and 40.8 N peak force. This is a reminder that “caught” and “caught softly” are different outcomes. Protocol and thresholds.
Final policies are trained with three seeds and evaluated on the same 3,000 held-out task conditions. At matched dataset size, success-only imitation reaches 83.8% task success and 41.1 N peak force. Refinement alone lowers force to 36.2 N but reaches 83.2% success. Repair alone reaches 88.5% success and 42.7 N. Combining repair and refinement reaches 90.3% and 37.3 N.
The full policy improves all ten shape-size combinations, with its largest reported gain—14.5 percentage points—on the large box. An ablation also shows that conditional geodesic search produces 44.3% repair, 59.8% refinement and 90.3% student success, versus 5.6%, 8.3% and 83.4% for direct action perturbations.
What this means for robotics
RoboSkin analysis: the contribution is a data-selection method for the moment when contact dynamics matter most. Rather than treating a full trajectory as uniformly valuable, it concentrates computation on the pre-contact segment that controls impact and grasp stability. That idea could complement tactile world models by identifying which action windows deserve denser sensing and validation.
The paper also shows why success-only data can be misleading for contact-rich work. A system can maximize completion by accepting large impact. For fragile objects or human-facing manipulation, force, relative velocity and follow-through should remain visible metrics beside task success.
Limitations and availability
Outcome-Sensitive Motion Search is an arXiv v1 preprint and a simulation-only result. Candidate selection depends on simulated contact dynamics and an error model calibrated from one imitation policy. The search cannot recover behavior outside the learned motion manifold, including failures that require different finger control or post-contact recovery.
No official project page, repository, dataset download, trained model or software license was verified. The manuscript is readable through arXiv, but the 99,328 demonstrations described in the experiment are not thereby a publicly downloadable dataset. RoboSkin.ai has not run the simulator or reproduced the numbers.


