TaRL learns contact-rich rewards from tactile demonstrations
TaRL regresses task progress from successful and failed tactile demonstrations, then uses that signal to shape contact-rich reinforcement learning.

Researchers from National Taiwan University and Delta Electronics released Tactile Reward Learning, or TaRL, on September 29, 2026. TaRL learns a dense progress signal from sequences of tactile deformation maps, then adds that signal to reinforcement learning for contact-rich tasks. The reported gains include nut-threading success rising from 34% to 56% in simulation and cube pickup rising from 37% to 97% on a physical robot. Paper and version record.
Key takeaways
- TaRL treats touch as reward supervision, not only as a policy observation. Successful, rewound and failed tactile sequences teach a causal model to estimate task progress.
- The simulation reward model uses 200 successful and 800 failed demonstrations per task. Real-world reward learning uses 40 successful and 40 failed trajectories per task, while the offline policy dataset contains another 90 successful and 90 failed trajectories.
- The real-world percentages are averaged across three random seeds, but the manuscript does not state the number of inference trials per seed. The 60-point cube-pickup gain therefore lacks a disclosed evaluation denominator.
What changed
Reward engineering is a persistent bottleneck in reinforcement learning. A sparse success flag gives little guidance before a task is complete, while a hand-written dense reward can favor the wrong behavior. Video-based reward models offer another route, but scene appearance does not reliably expose whether a grasp is firm or whether contact force is correctly directed.
TaRL replaces the video sequence with tactile deformation maps from two fingertips. A shared three-layer convolutional encoder processes each map, and a causal Transformer estimates progress using only the history available at that step. Successful demonstrations receive a target that rises with time; failed demonstrations receive zero. Rewound successful sequences teach the model that undoing progress should reduce the reward. Method details.
The learned value is a shaping reward. It supplements rather than replaces the task's basic sparse or stage reward, and it is trained separately for each task. In simulation, policies use Proximal Policy Optimization. The physical SO-101 arm experiments use Implicit Q-Learning on a fixed dataset, so their improvement is an offline-policy result rather than evidence of unrestricted online exploration.
Results under the reported conditions
The simulated suite covers box placement, peg insertion, gear assembly and nut threading with a Franka Panda. Demonstrations are collected at one object position, while policies and reward quality are evaluated at other positions. The most concrete final-success comparison is nut threading: 34% without TaRL and 56% with it, an increase of 22 percentage points. Held-out reward tests use 50 successful and 150 failed trajectories for both in-domain and out-of-distribution positions.
The authors also compare TaRL with ReWiND, a visual reward learner using the same high-level progress-regression recipe. Visual features shift when the object moves, while the tactile representation remains more similar across positions. Combining visual and tactile rewards improves learning further on the three compared tasks, supporting the narrower claim that the modalities contribute different signals.
On the physical robot, cube pickup and peg insertion each use 80 demonstrations for the reward model and 180 different trajectories for offline policy learning. The cube-pickup score rises from 37% to 97%. For peg insertion, the paper separates pickup and insertion: TaRL adds 45 and 10 percentage points respectively. Those results are reported across three training seeds, but no test-rollout count or confidence interval is supplied.
What this means for robotics
RoboSkin analysis: TaRL moves tactile sensing one step upstream in the learning stack. Instead of asking a policy to discover how touch relates to success, it first converts contact history into an explicit training signal. This could be useful when teams have tactile demonstrations but cannot write a trustworthy force-aware reward.
The method also exposes a scaling trade-off. Its reward model is local enough to tolerate position changes and, in one box-to-can experiment, a new object instance. Yet it still needs task-specific successful and failed touch sequences. It is not a general tactile foundation reward, and it does not eliminate data collection for a new behavior.
Limitations and availability
TaRL is an arXiv v1 preprint, and RoboSkin.ai has not reproduced its experiments. The comparison isolates modality carefully, but simulation and hardware use different RL algorithms. The physical study covers two tasks on one SO-101 setup, and the missing inference-trial denominator limits statistical interpretation of its largest percentage gain.
The official project page provides method explanations, plots and videos. At verification time it did not expose a public implementation repository, downloadable demonstrations, trained reward models or a software/data license. The arXiv manuscript's availability does not make those implementation assets reusable.


