TacEx makes tactile uncertainty a target for robot exploration
TacEx decomposes model uncertainty by sensing modality and rewards uncertainty in touch, steering robot learning toward contact-rich experience.

Researchers from ETH Zurich, the University of California, Berkeley and the University of Texas at Austin released TacEx on September 30, 2026. The framework separates a learned dynamics model's uncertainty into visual, tactile and latent-state components, then lets a robot prioritize uncertainty in touch. The reported experiments show more contact-rich reward-free exploration and more sample-efficient post-training of a frozen vision-language-action policy, but all tests are in simulation. Paper and version record.
Key takeaways
- TacEx changes the exploration objective, not merely the policy input: ensemble disagreement about future tactile observations becomes an intrinsic reward.
- Reward-free collection is evaluated across five random seeds. VLA post-training covers eight contact-rich LIBERO-90 tasks and ten seeds, with 500,000 environment steps for most tasks and one million for task 69.
- The paper reports curves rather than a single aggregate success claim. Tactile-plus-visual uncertainty gives the highest reported mean final return on each of the five tasks in the extended comparison, but the benefit varies by task.
What changed
Standard curiosity methods reward predictive uncertainty wherever it occurs. In manipulation, that can spend a large budget on novel motion through empty space. TacEx instead models uncertainty separately for visual features, tactile force maps and the fused latent state. Weighting the tactile term more heavily biases the policy toward states where contact dynamics are unknown. Method and experiments.
The first experiment removes task reward entirely. A simulated parallel-jaw gripper explores scenes with one object and then four objects. After collection, the replay buffer is frozen, relabelled with a sparse pick-and-place reward and used to train Soft Actor-Critic offline. Tactile-weighted variants interact with and attempt to grasp objects more often than variants without tactile disagreement. Adding visual disagreement spreads interactions over roughly two objects, while tactile-only curiosity concentrates closer to one.
The second experiment applies the same idea to a pretrained policy. A lightweight learner predicts noise for a frozen diffusion-based VLA; the base VLA still consumes its original image, proprioception and language inputs. TacEx gives the noise-steering learner tactile observations and an intrinsic disagreement bonus. This is post-training around a frozen policy, not tactile pretraining of the VLA itself.
Results under the reported conditions
The LIBERO study uses task IDs 58, 69, 61, 77, 55, 21, 56 and 20. The simulated Panda gripper receives 32 × 64 × 3 tactile force maps and 64 × 64 images. Every main comparison uses the same pretrained policy checkpoint and per-task interaction budget. The authors report mean evaluation return with standard-error bands across ten seeds.
Five tasks receive the most extensive baseline comparison: ketchup, soup and butter into a tray; a book into the back of a caddy; and turning on a stove before placing a frying pan. Tactile-plus-visual disagreement has the highest reported mean final return on all five. Additional controls aggregate the taxels into one signal, replace SAC with PPO or update the diffusion action head. The spatial tactile-plus-visual variant remains strongest under the displayed protocol.
RoboSkin verification: the paper does not publish one numerical average across all eight tasks, so this article does not manufacture one from plotted curves. It also distinguishes tactile input from tactile-directed exploration. The ablation shows that adding touch to the learner state helps, while the best variants additionally predict spatial tactile outcomes as the disagreement target.
What this means for tactile robotics
RoboSkin analysis: TacEx treats a tactile sensor as a guide to where the robot should collect experience. That is different from using touch only after contact has already occurred. For contact-rich manipulation, it could make data collection more intentional by spending fewer transitions away from objects.
The multi-object result also exposes a useful design trade-off. Tactile-only curiosity can over-focus on one reliable source of contact; combining vision and touch maintains broader scene coverage. A deployment may therefore need adaptive modality weights rather than one permanent setting. That question links exploration design to the wider robot-learning pipeline, not just sensor choice.
Limitations and availability
TacEx is an arXiv v1 preprint marked under review, and RoboSkin.ai has not reproduced it. The study uses simulated force maps, one parallel-jaw gripper and no fragile or deformable objects. It does not test tactile noise, calibration drift, sensor wear, safe-force constraints or physical resets. Contact-seeking curiosity by itself does not prevent unsafe contact.
The authors state that the representation could accept vision-based tactile images, but no multi-finger hand or real sensor is evaluated. At verification time, neither the arXiv record nor manuscript linked a public project page, code repository, dataset, checkpoint or implementation license. The article is CC BY 4.0; that license applies to the paper, not to absent software or data.


