Structured benchmark directory
Tactile robotics benchmarks compared
Compare tactile robotics benchmarks by task, sensor, robot, modality, metric, split protocol, access, and evidence boundary.
Published 2026-08-19 | Updated 2026-08-22 by Steven Yang

- 4
- sections
- 4
- questions
- 7
- next routes
Short answer
What you need to know
- 1
A tactile robotics benchmark is a defined evaluation contract: task, sensor input, robot or collection platform, data split, metric, baseline, and version. A dataset becomes a benchmark only when it is paired with a reproducible protocol.
- 2
Benchmark scores are not portable across different tactile sensors, robots, objects, control rates, or success definitions. Compare systems only inside the same protocol or reproduce both under one shared setup.
- 3
For manipulation, always include an outcome-level baseline with touch disabled. Representation accuracy alone does not prove that tactile sensing improves robot behavior.
Topic 01
What makes a tactile benchmark credible?
A credible benchmark states what is held constant and what is allowed to vary. At minimum, it identifies the physical data unit, sensor and robot configuration, training information, split unit, metric, baseline, evaluation repetitions, and failure definition.
Tactile data are unusually vulnerable to leakage because adjacent frames from the same press, grasp, or trajectory can be nearly identical. The correct split unit follows the claim: hold out presses for new-contact evaluation, materials for material generalization, sensors for cross-sensor transfer, and objects or tasks for manipulation transfer.
- Input contract: raw tactile images, taxel arrays, force vectors, touch plus vision, or learned embeddings
- Independence unit: frame, contact sequence, object, material, task, sensor, robot, or laboratory
- Outcome: perception accuracy, prediction error, task success, recovery, latency, or safety event
- Reproducibility: released data, split manifest, evaluation code, hardware description, and version
Topic 02
Benchmark families are not one leaderboard
The directory includes representation suites, multimodal understanding tests, cross-sensor transfer protocols, full-hand benchmarks, and closed-loop manipulation challenges. They answer different questions and should not be collapsed into a single ranking.
| Benchmark family | Primary question | Typical unit | Useful output | Common mistake |
|---|---|---|---|---|
| Representation | Does an encoder preserve contact information that transfers to downstream tasks? | Tactile frame, clip, or contact sequence | Frozen-encoder task metrics and data-efficiency curves | Treating average improvement across unlike tasks as one universal score |
| Cross-sensor | Does a model survive a change in sensor design or instance? | Held-out sensor or sensor instance | Zero-shot gap and few-shot recovery | Mixing frames from the same physical contact across train and test |
| Multimodal | Can touch align with vision, language, audio, or robot state? | Aligned example or trajectory | Retrieval, classification, generation, or prediction metrics | Ignoring timestamp error or generated-label quality |
| Manipulation | Does tactile input improve closed-loop task behavior? | Episode or repeated real-robot trial | Success, recovery, time, damage, and intervention rate | Reporting perception accuracy without a touch-disabled task baseline |
Topic 03
Minimum evaluation contract for tactile manipulation
A useful task-level comparison fixes the gripper or hand, object set, initial-state distribution, controller budget, tactile preprocessing, vision inputs, and success rule. It then repeats trials across relevant disturbances and reports both aggregate results and failure categories.
- No-touch or disabled-touch baseline on the same robot and task
- Repeated trials across objects, poses, contact conditions, and disturbances
- Latency from physical contact through sensing, inference, command, and actuation
- Calibration drift, sensor replacement, wear, and reinitialization procedure
- Failure taxonomy: miss, slip, jam, drop, excess force, timeout, or unsafe contact
Topic 04
How to use the directory
Filter by benchmark type, task, sensor, and year, then open the primary paper and official project or code page. Record the exact version you use. “Open source” on a project page does not automatically establish the license of every dataset, model weight, or bundled asset.
The evidence-boundary column is deliberate. It prevents a benchmark’s strongest reported result from being generalized beyond its sensor, robot, task, split, or publication status.
Database method
How the tactile benchmark directory is built
This directory separates named benchmarks and author-defined evaluation protocols from generic dataset scale or isolated headline scores. Every record identifies tasks, modalities, sensors, robots, metrics, protocol, access, and a limitation boundary.
- Records
- 12
- Reviewed through
- 2026-08-22
Inclusion rule
- The primary source must document a repeatable task and evaluation protocol, not only a demonstration.
- Reported metrics remain attached to the named robot, sensor, task set, baselines, and rollout conditions.
Editorial normalization
- Percentage points, averages, rollout counts, and task names preserve the source protocol.
- Simulation, physical-robot, author-run, and independently hosted evidence are labeled separately.
Excluded claims
- No score is promoted to a universal leaderboard when protocols, sensors, robots, or splits differ.
- A project demo or dataset size is not treated as benchmark performance.
Known limitations
- Most current tactile benchmarks remain author-defined and cannot be ranked across incompatible protocols.
- Directory inclusion does not establish independent reproduction, production safety, or statistical significance.
Structured benchmark explorer
Compare tactile robotics benchmarks
These suites evaluate different sensors, modalities, robots, and outcomes. A score is comparable only when the task, split, hardware, metric, and benchmark version match.
Source review: 2026-08-22 / 12 records
Showing 12 of 12 benchmarks
| Benchmark | Type / year | Tasks | Modalities / hardware | Metrics / protocol | Access / evidence boundary | Primary links |
|---|---|---|---|---|---|---|
| UniVTAC BenchmarkScaleLab, Shanghai Jiao Tong University; D-Robotics; ViTai Robotics; The University of Hong Kong; Nanjing University; Shenzhen University; Wuhan University; Fudan University; Tsinghua UniversityReviewed 2026-08-22 | Simulation-based visuo-tactile manipulation benchmark 2026 | Lift Bottle; Pull-out Key; Lift Can; Put Bottle in Shelf; Insert Hole; Insert HDMI; Insert Tube; Grasp Classify | Signals: Head RGB; Wrist RGB; Robot state and action; Bilateral simulated tactile RGB; Marker motion; Gelpad depth; Tactile-sensor pose Sensors: Bilateral simulated GelSight Mini sensors in the released collection and evaluation pipeline Robot: Simulated Franka Panda with a parallel-jaw gripper | Binary task success rate over evaluation rollouts; Per-task success rate; Eight-task macro-average success rate Protocol: The paper trains every policy on 50 automatically collected full trajectories per task, or 400 trajectories across eight tasks, and evaluates each method on 100 test rollouts per task. Table I reports eight-task averages of 30.9% for vision-only ACT, 40.5% for VITaL, and 48.0% for ACT with the UniVTAC Encoder; 30.9% to 48.0% is an increase of 17.1 percentage points. | The paper, project page, Apache-2.0 repository, collection and evaluation code, and a separately MIT-labeled Hugging Face package of 800 HDF5 episodes are public. The current repository collection and evaluation workflow supports simulated GelSight Mini; ViTai GF225 and Xense WS are listed as planned additions. Boundary: This is an author-defined simulation benchmark, not an independent leaderboard or physical-robot safety test. The public 800 episodes are not the paper’s 400 policy-training trajectories or its evaluation rollouts. Hosted checkpoint logs reviewed separately do not reproduce the exact Table I averages, so the paper and hosted logs must not be presented as identical runs. | |
| TouchWorld Real-Robot Evaluation ProtocolHarbin Institute of Technology, Shenzhen; PHANES AIReviewed 2026-08-22 | Author-defined real-robot manipulation evaluation protocol 2026 | Water Flower; Tabletop Clearing; Cup Insertion; Power Plug Insertion; Pot Wiping; Tissue Pulling | Signals: Multi-view RGB; Proprioceptive state; Left and right tactile pressure observations; Natural-language task instruction; Robot actions Sensors: JQ-Industries tactile glove mounted on Wuji dexterous hands Robot: One unnamed humanoid platform with Wuji dexterous hands | Per-task manipulation success rate in a clean setting; Per-task manipulation success rate under human perturbation; Six-task average success rate Protocol: The v2 preprint defines six contact-rich tasks, collects 200 teleoperated training trajectories per task, and conducts 100 real-robot evaluation rollouts per task. Each task is evaluated in clean and human-perturbation settings; the paper does not state that this is an independent shared leaderboard. | The preprint and official project page are public. No standalone benchmark package, evaluation scripts, fixed test-set release, or benchmark artifact license was verified on 2026-08-22. Boundary: This is an author-defined evaluation on one unnamed humanoid, one Wuji-hand and tactile-glove configuration, six tasks, and source-specific baselines. The reported 65.0% clean and 53.7% perturbed averages are author-reported TouchWorld results, not independent validation or a universal cross-model benchmark. | |
| SoftVTBenchTuojing Intelligence; Tsinghua University; King's College London; Southeast University; Stevens Institute of Technology; Hong Kong University of Science and Technology (Guangzhou); University of Manchester; Simple AI; Imperial College London; Carnegie Mellon University; Zhejiang University; Beihang University; University of Hong KongReviewed 2026-08-22 | Simulated deformation-aware visuo-tactile manipulation benchmark 2026 | Object-Soft manipulation; Spatial-Soft manipulation; Object-Rigid matched control; Spatial-Rigid matched control; In-distribution evaluation; Out-of-distribution evaluation | Signals: Simulated multi-view RGB; Simulated bilateral tactile RGB; Simulated marker motion; Proprioception; Language; Robot actions; Evaluator-only FEM deformation state Sensors: Simulated GelSight Mini profiles through TacEx, Taxim, and FOTS Robot: Simulated Franka arm with Panda parallel-jaw gripper | Task Success Rate (TSR); Deformation-aware Success Rate (DSR); Peak normalized deformation; TSR-DSR gap Protocol: Uses policy-independent, object-specific probing to fix deformation tolerances before training; DSR credits a rollout only when it completes the task and remains within the calibrated FEM deformation tolerance. | The official project and GitHub repository link public training and evaluation code plus a Hugging Face dataset mirror. Hugging Face dataset-card revision fd2793a documents 4,000 hosted demonstrations, but the older GitHub README still lists 1,628 demonstrations and 33 assets; users should pin a revision because the first-party release documents disagree. Boundary: All robot, tactile, object, and FEM signals are simulated; no SoftVTBench-specific physical-sensor or sim-to-real validation is reported. DSR is defined by this paper and is not an industry standard, safety certification, or universal damage metric. | |
| HT-BenchBeihang University; Rimbot; ShanghaiTech University; Tsinghua University; Chinese Academy of Sciences; BUPTReviewed 2026-08-22 | Full-hand representation benchmark 2026 | Tactile similarity retrieval; Masked tactile inpainting; Vision-to-tactile synthesis; Tactile frame prediction; Downstream board cleaning; Downstream pear picking; Downstream water pouring; Downstream sand shoveling | Signals: Egocentric RGB; Full-hand tactile pressure maps Sensors: Single reported full-hand tactile sensing pipeline Robot: Dexterous full-hand platform | Recall@5; Full-map and hole-region RMSE; Full-map and hole-region cIoU; Real-world task success rate Protocol: Four offline tracks use standard-test and task-level OOD splits. Version 2 also reports board cleaning, pear picking, water pouring, and sand shoveling with 15 trials per method for each task. | The v2 preprint says the authors will release all data, evaluation protocols, pretrained weights, and training/testing scripts; no dedicated downloadable package or official repository was verified on 2026-08-22. Boundary: A 2026 preprint tied to one full-hand tactile system and embodiment pipeline; it does not cover fingertip optical tactile sensors, force/torque sensors, skin-like taxel arrays, or non-hand embodiments. The four-task downstream table reports success rates without confidence intervals or statistical-significance testing and is not a universal cross-sensor or cross-robot standard. | |
| TacBenchFAIR at Meta; University of Washington; Carnegie Mellon UniversityReviewed 2026-08-19 | Tactile representation evaluation suite 2024 | Force estimation; Slip detection; Pose estimation; Grasp stability; Textile recognition; Dexterous manipulation | Signals: Vision-based tactile images Sensors: DIGIT; GelSight 2017; GelSight Mini Robot: Task-specific manipulation platforms | Task-specific regression; Task-specific classification; Manipulation outcome Protocol: Frozen tactile encoders are paired with lightweight downstream decoders across six touch-centric tasks. | Official project page links the paper, code, datasets, and model weights. Boundary: The suite targets vision-based tactile sensing; average gains across different task metrics should not be read as one universal score. | |
| ObjectFolder BenchmarkStanford University; Carnegie Mellon UniversityReviewed 2026-08-19 | Multisensory object benchmark suite 2023 | Cross-sensory retrieval; Contact localization; Material classification; 3D shape reconstruction; Visuo-tactile cross-generation; Grasp stability; Contact refinement; Surface traversal; Dynamic pushing | Signals: Vision; Audio; Touch; 3D geometry Sensors: GelSight robotic finger for ObjectFolder Real Robot: Franka Emika Panda for ObjectFolder Real collection; Simulated object interactions | Task-specific retrieval; Classification; Reconstruction; Manipulation metrics Protocol: Ten tasks are organized around recognition, reconstruction, and manipulation with neural and real objects. | The official project page links open datasets, papers, downloads, and the benchmark suite. Boundary: Results across simulated neural objects and real-object measurements are not interchangeable; each task keeps its own protocol. | |
| Touch-Vision-Language BenchmarkUC Berkeley; Meta AI Research; TU Dresden; CeTIReviewed 2026-08-19 | Multimodal tactile-language benchmark 2024 | Tactile-semantic description; Touch-vision alignment; Touch-language alignment | Signals: DIGIT tactile image; RGB image; English tactile description Sensors: DIGIT Robot: Handheld collection device; Robot-collected SSVTP subset | Tactile-semantic classification accuracy; Cross-modal alignment accuracy; LLM-rated description consistency Protocol: TVL combines HCT and SSVTP data and evaluates multimodal models on tactile-semantic understanding and alignment. | The official project page links the paper, code, dataset, and models. Boundary: Most HCT language labels are VLM-generated; the benchmark paper documents label errors and should be read with that boundary. | |
| TacVerseResearch team listed in the primary preprintReviewed 2026-08-19 | Cross-sensor vision-based tactile benchmark 2026 | Shape classification; Grating classification; Force regression | Signals: Vision-based tactile images; Force labels Sensors: Seven vision-based tactile sensor designs Robot: Controlled tactile data-collection platform | Classification accuracy; Force-regression error; Cross-sensor transfer change Protocol: Compares within-sensor training, zero-shot cross-sensor transfer, and few-shot adaptation on 106,800 tactile images. | The preprint describes the dataset and protocol; no separate official download page was verified. Boundary: A 2026 preprint. Cross-sensor results are specific to the seven sensors, three tasks, and released split design. | |
| RCT Generalization BenchmarkTU Dresden; ScaDS.AI Dresden/Leipzig; LASR LabReviewed 2026-08-19 | Contact-sequence and material generalization protocol 2026 | Touch-to-text retrieval; Touch-to-vision retrieval; Material and category probes; Cross-sensor generalization | Signals: DIGIT tactile images; Material RGB; Language; Normal force; Indentation depth Sensors: Three DIGIT sensors Robot: Robot arm with rotating sensor adapter | Recall@1; Category-probe accuracy; Generalization gap Protocol: Preserves full robot presses and supports held-out sequences, contact positions, sensors, categories, and materials. | Official project page links the dataset, split tools, code, and preprint. Boundary: The benchmark focuses on industrial reference materials and contact sequences rather than closed-loop manipulation success. | |
| ManiSkill-ViTac 2025Challenge organizing team listed in the primary paperReviewed 2026-08-19 | Simulation-to-real manipulation challenge 2025 | Tactile manipulation; Vision-tactile fusion manipulation; Tactile sensor structure design | Signals: Vision; Marker-based visuo-tactile sensing; Robot state Sensors: Simulated and real marker-based visuo-tactile sensors; GelSight Mini integration resources Robot: Challenge simulation and real-robot environments | Standardized task metrics defined per challenge track Protocol: Three independent tracks evaluate learned contact-rich manipulation and sensor design in simulated and real-world settings. | The official repository provides environments, leaderboard links, and real-robot resources under Apache-2.0. Boundary: Challenge scores depend on a specific annual track, environment version, task definition, and evaluation rules. | |
| TactiDexShanghaiTech University; InstAdaptReviewed 2026-08-19 | Real-world tactile-guided dexterous manipulation benchmark 2026 | Single-hand dexterous manipulation; Bimanual manipulation; Human-to-robot skill transfer | Signals: Whole-hand pressure; Hand kinematics; Object 6D pose; Language; Task phase Sensors: Whole-hand tactile glove Robot: Bimanual Franka Inspire platform; Human demonstration capture | Standardized task and transfer metrics defined by the source Protocol: Aligns whole-hand tactile signals with multi-granularity kinematic and object states for standardized real-world evaluation. | The official project page documents the benchmark and demonstrations; a separate public download was not verified. Boundary: A 2026 preprint; inspect the current release, task definitions, splits, and access terms before comparison. | |
| VTDexManipZhejiang UniversityReviewed 2026-08-19 | Visual-tactile dexterous policy benchmark 2025 | Six dexterous manipulation tasks; Cross-task generalization; Modality ablation | Signals: Vision; Sparse binary touch; Proprioception Sensors: Sparse binary tactile sensing on the manipulation platform Robot: Dexterous manipulation platform described by the authors | Task success; Generalization; Noise and viewpoint robustness Protocol: Compares pretrained and non-pretrained visual-tactile policies across six tasks using a dataset spanning 10 daily activities and 182 objects. | The author project page links the ICLR paper plus code and dataset. Boundary: Reported policy gains are specific to the paper platform, sparse tactile representation, task suite, and RL training setup. |
Common questions
FAQ for this topic
What is a tactile robotics benchmark?
It is a reproducible evaluation contract for a tactile perception, representation, or robot-control question. It specifies inputs, hardware, data splits, metrics, baselines, and evaluation conditions.
Can benchmark scores be compared across tactile sensors?
Only when both systems use a shared protocol that controls sensor mounting, calibration, data, robot, task, metric, and evaluation version. Otherwise the numbers describe different experiments.
Is a tactile dataset automatically a benchmark?
No. A dataset supplies observations. A benchmark adds defined tasks, splits, metrics, baselines, and evaluation code or instructions.
Which baseline matters most for tactile manipulation?
Use the same robot, task, controller budget, and visual inputs with the tactile pathway disabled. That isolates whether touch improves the task outcome.
