TACTIC finds no single best tactile encoder for contact-rich manipulation
TACTIC crosses five tactile encoders with five conditioning methods on four real robot tasks. Its central result is conditional: the best combination changes with the contact problem.

Researchers led by the University of Technology Nuremberg released TACTIC on September 25, 2026. TACTIC is a real-robot study of how five tactile encoders and five policy-conditioning methods interact across wiping, USB-C insertion and screw insertion. After 2,180 evaluation rollouts, the answer is not a universal winner: encoder and fusion choices depend on the task. Paper and version record.
Key takeaways
- TACTIC evaluates 25 encoder-conditioning combinations on four tasks, with 20 rollouts for each reported ablation condition.
- The simplest ResNet18 plus concatenation reaches 80% termination on USB-C insertion, compared with 55% for the vision-only policy in that task.
- Four official datasets are downloadable now and contain 410 episodes in total. The project page still labels policy code “coming soon.”
What changed
Many tactile-policy papers change the sensor representation and fusion mechanism at the same time, making the source of an improvement hard to identify. TACTIC instead fixes an Action Chunking with Transformers policy and crosses five encoders—ResNet18, SARL, UniT, Sparsh and T3—with concatenation, FiLM, gated cross-attention, CLIP-style replacement and CLIP-style token conditioning.
The hardware uses two Franka FR3 arms in a leader-follower collection setup. The follower has two DIGIT sensors on a Franka Hand, an Intel RealSense D405 at the wrist and a D435 side camera. The team collects about 100 demonstrations per task at 30 Hz, using four pairs of replaceable DIGIT cartridges to introduce sensor variation. The policy predicts 120-step action chunks. Experimental setup.
What the real-robot study found
USB-C insertion favors the uncomplicated baseline: ResNet18 with feature concatenation reaches 80% termination, while vision-only reaches 55%. Screw insertion instead favors semantic alignment, with Sparsh-CLIP-T and T3-CLIP-T each reaching 90% termination under the reported protocol. On other tasks, the same CLIP-style mechanisms can reduce performance. That reversal is the paper's main engineering result.
The authors also introduce out-of-distribution “disco” lighting. Non-CLIP variants fail under that visual shift, while some Sparsh-plus-CLIP configurations preserve partial performance. Each out-of-distribution condition uses ten trials, so these percentages have wider uncertainty than the 20-rollout main ablations. Results and OOD analysis.
The headline count needs careful interpretation. The 2,180 figure is evaluation rollouts across policies and conditions, not the size of the imitation-learning dataset. The official Hugging Face collection currently lists four datasets—screw, board, vase and USB-C—with 104, 102, 101 and 103 episodes respectively, or 410 episodes total. Dataset cards display an MIT license, 30 fps data and six cameras; several README bodies are still empty, so documentation depth does not yet match file availability.
What this means for robotics
RoboSkin analysis: “use a pretrained tactile encoder” is not enough of a recipe. The representation must preserve the contact information a task needs, and the policy must expose that information in a useful way. A semantic alignment layer may help when tactile and visual observations share a stable task concept, but it can also discard continuous detail needed for surface interaction.
For teams building visuo-tactile policies, TACTIC argues for a small task-specific ablation before scaling data collection. A simpler encoder can outperform a larger pretrained model, especially when latency, cartridge changes and cross-device variation matter more than broad representation learning.
Limitations and availability
TACTIC is an arXiv v1 preprint, and RoboSkin.ai has not reproduced its 2,180 rollouts. The study uses one robot family, DIGIT sensors and four tasks. There is no confidence interval around most reported success rates, and 20 trials per setting cannot resolve small differences reliably. Performance under new sensor geometries or policies remains unknown.
The project page marks code as “coming soon.” The four datasets are publicly downloadable and display MIT licenses, but users should inspect each repository and file manifest before reuse. The arXiv manuscript uses arXiv's submission license; that does not license unreleased training code.


