<- Back to news

Wrench-ACT makes force and torque the robot policy action

Wrench-ACT pairs force-reflecting bilateral demonstrations with an ACT policy that directly commands a six-dimensional target wrench.

Preprint · arXiv v1 · five UR5e tasks · 50 rollouts per task-policy pair · demonstration release promised, not verified publicSource date: Read the primary source ↗
direct wrench controlforce-feedback teleoperationcontact-rich manipulationimitation learning
Diagram of bilateral force-feedback demonstrations training an ACT policy that outputs a six-dimensional wrench to a force controller.
Original RoboSkin.ai schematic of Wrench-ACT data collection and control. It is explanatory artwork, not an experimental figure.

Researchers from Siemens and the University of Technology Nuremberg released Wrench-ACT on September 29, 2026. The system modifies the usual imitation-learning contract: instead of predicting a target position and letting controller error create contact force, its Action Chunking with Transformers policy outputs a six-dimensional force/torque wrench directly. With matching force-reflecting demonstrations, it averages 76.8% success across five contact-rich tasks versus 40% for a position-policy baseline. Paper and version record.

Key takeaways

  • Direct wrench output works best only when the demonstrations contain intentional wrench commands. Converting ordinary position demonstrations into approximate wrenches averages 28.8% in the same action space.
  • The main table evaluates four collection/action combinations on five tasks with 50 rollouts each: 1,000 physical evaluation rollouts in total.
  • The paper promises more than 1,000 wrench-action demonstrations on a companion site upon publication, but no public companion URL, archive or license was verified with the v1 release.

What changed

Position policies command where the tool should go. During contact, a Cartesian impedance controller turns the pose error into force. That is useful but indirect: two identical target poses can produce very different interaction loads as geometry and stiffness change.

Wrench-ACT removes the pose target from the learned action. Three RGB streams, robot state and gripper position enter a standard single-task ACT model. Its seven-dimensional output contains a six-axis target wrench plus the gripper command. A UR5e force controller executes the wrench at 500 Hz while policy inference runs near 50 Hz and holds the latest target between predictions. Method and controller.

Data collection is the other half of the design. In the bilateral setup, a human pushes a leader arm and feels the follower's reaction. The measured leader wrench becomes the action label. A second dataset uses a Meta Quest position interface and Cartesian impedance control. The authors also translate each dataset into the opposite action representation, producing four policy conditions that separate the collection interface from the learned output.

Results and the matched-data effect

The five tasks are peg insertion, fuse clipping, fan insertion, industrial-connector mating and pen writing. Bilateral data plus wrench output records 74%, 88%, 58%, 80% and 84%, averaging 76.8%. The conventional VR-data/position-action condition records 6%, 96%, 8%, 12% and 78%, averaging 40%.

That headline average hides two qualifications. Wrench-ACT is lower on fuse clipping, 88% versus 96%, and only three task differences are statistically significant after the paper's multiple-comparison correction: peg, fan and industrial-connector insertion. Pen writing and fuse clipping do not establish a significant advantage.

The cross-condition table is more revealing. A wrench policy trained on wrenches reconstructed from VR position trajectories averages 28.8%. A position policy derived from bilateral wrench data averages 34.8%. The strongest outcome appears when the interface and action agree: humans deliberately command force, and the policy predicts the same quantity.

An inference-rate ablation on the industrial connector reports 80% at 50 Hz, 36% at 30 Hz, 20% at 15 and 5 Hz, and 12% at 1 Hz. The nominal result uses 50 trials; each reduced-rate condition uses 25. When the paper support under the pen is raised 2.5 centimeters, the bilateral-wrench policy falls from 84% to 76%, while the VR-position policy falls from 78% to 24%.

What this means for robotics

RoboSkin analysis: Wrench-ACT is evidence for aligning a teleoperation interface with the quantity a policy must control. Merely logging a force/torque sensor alongside position commands does not mean the dataset contains deliberate force strategy. For contact-rich data, action provenance matters as much as the presence of force channels.

Direct wrench control is not a universal replacement for pose actions. The authors deliberately choose tasks where interaction force is central, and they state that incidental-contact tasks may not benefit. The system also depends on a capable inner force loop and a mechanically compliant setup; the learned policy is only one layer of the control stack.

Limitations and availability

Wrench-ACT is an arXiv v1 preprint, and RoboSkin.ai has not reproduced it. All experiments use one UR5e-based setup and single-task models trained from scratch. Datasets contain 200–300 episodes per task and interface, but operator count is not stated. The force-reflecting collection rig is substantially more specialized than common handheld or VR capture systems.

The manuscript says a dataset of more than 1,000 wrench-action demonstrations will be released on a companion website “upon publication.” At verification time the arXiv record did not link that site, code, downloadable data, checkpoints or a release license. A promised release should not be treated as an available dataset.

Continue the topic

Impact-aware manipulationOutcome-sensitive search teaches a dexterous hand to catch with less impactDexterous policy learningDexTaG uses human touch to guide dexterous policy retargetingVisuo-tactile learningVisTacAlign turns human touch into robot training data