<- Back to news

UVTA scales dexterous policies with human tactile demonstrations

UVTA aligns human and robot touch-action trajectories, combining 1,000 human and 150 robot demonstrations per task to train one contact-aware dexterous policy.

Preprint · arXiv v1 · five real-robot tasks · project page public · linked code repository unavailable at verificationSource date: Read the primary source ↗
human tactile demonstrationsvisual tactile action modelingdexterous manipulationcross-embodiment learning
Diagram of human tactile demonstrations and robot demonstrations entering a shared visual tactile action model.
Original RoboSkin.ai schematic of UVTA cross-embodiment co-training. It is explanatory artwork, not an experimental figure.

Researchers from Shanghai Jiao Tong University, the Beijing Academy of Artificial Intelligence, Sharpa Robotics and Beijing Institute of Technology released Unified Visual-Tactile-Action Modeling, or UVTA, on September 28, 2026. The system co-trains on human and robot interaction data so that large volumes of human touch can improve a robot policy without treating human joint motion as directly executable robot commands. Paper and version record.

Key takeaways

  • Each of five tasks uses 1,000 human demonstrations and 150 robot demonstrations, producing 5,750 demonstrations across the study.
  • The full model averages 70% stage-wise success over 10 physical trials per task, compared with 29% for the strongest reported baseline and 42% without future-tactile prediction.
  • The project links a GitHub repository, but that URL returned a public 404 during verification. No downloadable dataset, model weights or usable software license was confirmed.

What changed

UVTA starts with a wearable capture system: a passive 22-degree-of-freedom hand exoskeleton records finger motion, two VIVE trackers recover wrist pose, a wrist camera captures RGB and four force-sensing regions on each fingertip produce a 20-dimensional tactile vector. The authors map these human observations into the same numerical representation used by a 22-degree-of-freedom Sharpa Wave robot hand.

The model turns the current wrist image into a 384-dimensional visual token and the tactile input into a 20-dimensional token. Their concatenated 404-dimensional representation feeds a diffusion action head and a tactile regression head. Both human and robot sequences supervise the shared representation, but deployment decodes only robot actions. Architecture and dataset.

This is co-training rather than direct teleoperation transfer. Human hand actions still provide a learning signal, while embodiment-specific normalization and robot demonstrations teach the executable action space. The action head predicts 16 steps and executes eight before replanning from new visual and tactile observations.

What the real-robot evaluation found

The five tasks are page flipping, bayonet light-bulb insertion, toggle-switch operation, tactile ball classification and liquid transfer. Every method receives the same wrist RGB, fingertip touch and right-arm/right-hand action space, and each result averages ten real-robot trials with randomized object positions or heights.

UVTA reports an unweighted average of 70%: 83% on page flipping, 80% on the light bulb, 88% on the toggle switch, 56% on ball classification and 44% on liquid transfer. The strongest baseline, T-Rex, averages 29%; RDP averages 14%, and ViTacFormer averages 1%. Removing the future-tactile head but retaining touch reduces the average to 42%, while the vision-only ablation reaches 26%. Full comparison.

These are stage-wise scores, not purely binary task completion. For example, page turning assigns partial credit for separating pages or turning multiple pages before awarding full credit for turning exactly one. That scoring makes intermediate progress visible, but it also means the headline percentages should not be interpreted as 35 complete successes out of 50 trials.

The scaling experiment holds 150 robot demonstrations per task fixed and evaluates three tasks as human data grows. The average moves from 0% with no human trajectories to 8%, 29%, 51% and 84% with 250, 500, 750 and 1,000 human trajectories. The largest reported step occurs between 750 and 1,000, but only three of the five tasks are included in this scaling curve.

What this means for robotics

RoboSkin analysis: UVTA treats human touch as representation supervision rather than pretending away the embodiment gap. That is valuable for robot datasets, because people can generate diverse contact sequences much faster than they can teleoperate a high-dimensional robot hand.

The study also shows why dataset totals need context. The 5,000 human demonstrations provide variety, while 750 robot demonstrations anchor executable behavior. Human data did not replace robot data in this experiment; it complemented it. The method therefore supports a hybrid scaling strategy, not a claim that tactile robot demonstrations are unnecessary.

Limitations and availability

UVTA is an arXiv v1 preprint, and RoboSkin.ai has not reproduced it. Evaluation uses one fixed-base North robot, one Sharpa Wave embodiment and five tasks. Touch is measured only at four regions per fingertip, so palm contact, shear and distributed whole-hand pressure are not represented. Ten trials per task also make small score differences uncertain.

The project page provides demos and limitations. Its Code button points to github.com/uni-vta/UVTA, but that repository returned 404 through both the public page and GitHub API at verification time. The page describes the dataset but does not expose a downloadable archive. Public links therefore should not be equated with released training assets.

Continue the topic

Visuo-tactile learningVisTacAlign turns human touch into robot training dataVisuo-tactile dexterityOccluDex fuses 3D geometry and touch when robot hands block the viewRobot demonstration dataTouch2Robot adds simulated robot contact to human demonstrations