<- Back to research

HT-Bench full-hand tactile benchmark for robot manipulation

HT-Bench v2 pairs egocentric vision with millions of full-hand tactile frames, corrects the vision-to-tactile metric split, and adds four real-robot evaluations.

HT-Benchfull-hand tactile sensingegocentric visiontactile representation learning
Illustration for HT-Bench full-hand tactile benchmark for robot manipulation

Updated technical brief - August 22, 2026 (arXiv v2)

HT-Bench is an arXiv preprint benchmark for learning and evaluating dexterous full-hand tactile representations alongside egocentric vision. Version 2, submitted on August 20, 2026, reports 10 million RGB frames and 7.8 million tactile frames collected across 226 tasks. For robot skin, its value is a concrete evaluation structure: test contact geometry, cross-modal alignment, and bounded downstream transfer instead of treating tactile frame count as proof of useful robot learning.

Source findings

The HT-Bench paper starts from a real benchmarking problem: tactile sensors, data formats, and robot embodiments vary too much for one universal leaderboard. The authors therefore study a narrower route based on egocentric vision paired with full-hand tactile observations.

The source describes four benchmark tasks: fine-grained tactile similarity retrieval, masked tactile inpainting, vision-to-tactile synthesis, and multimodal tactile frame prediction. Together they test whether a representation preserves contact structure, aligns touch with vision, and generalizes beyond the tasks used for training.

The paper also introduces HandTouch, a vector-quantized vision-tactile encoder. In Table 2, HandTouch improves Recall@5 for fine-grained tactile retrieval from the strongest ViT baseline's 74.65% to 85.23%. For vision-to-tactile synthesis, the standard-test full-map cIoU is 0.628 for the ViT baseline and 0.689 for HandTouch; on the task-level OOD split, the corresponding values are 0.446 and 0.457. The standard-test and OOD columns must not be conflated.

For masked tactile inpainting on the standard test split, Table 2 reports HandTouch full-map RMSE of 0.009 and full-map cIoU of 0.912, compared with 0.022 and 0.762 for the ViT baseline. Section 5.1 instead states 0.010 and 0.911 for HandTouch, a small internal inconsistency in the preprint. This brief uses the Table 2 values as the primary quantitative record and preserves the discrepancy rather than silently mixing the two statements.

RoboSkin analysis

HT-Bench evaluates the representation layer, not whether one tactile sensor is universally best. That distinction matters because a hardware comparison and a learned-representation comparison answer different engineering questions.

Benchmark taskCapability testedRobot-skin questionEvidence boundary
Tactile similarity retrievalFine-grained contact representationDo related full-hand contact states remain close in the embedding?Retrieval does not prove closed-loop manipulation success.
Masked tactile inpaintingSpatial contact structureCan missing parts of a tactile observation be reconstructed from context?Reconstruction quality depends on the sensing layout and training distribution.
Vision-to-tactile synthesisCross-modal alignmentCan egocentric vision constrain likely contact patterns?Predicted touch is not a substitute for measured contact during deployment.
Multimodal frame predictionTemporal and cross-modal dynamicsDoes the model preserve how visual and tactile state change together?Frame prediction does not by itself establish safe robot control.

v2 real-world downstream evaluation

Version 2 adds four contact-rich real-robot tasks: board cleaning, pear picking, water pouring, and sand shoveling. The paper reports 15 trials per method for each task.

MethodBoard cleaningPear pickingWater pouringSand shovelingMean
ResNet-Scratch20.0%53.3%13.3%20.0%26.7%
ResNet-Trained46.7%73.3%33.3%46.7%50.0%
ViT40.0%73.3%33.3%33.3%45.0%
HandTouch66.7%86.7%53.3%66.7%68.3%

Across these four tasks, the author-reported HandTouch mean is 68.3%, versus 50.0% for the strongest baseline by mean, ResNet-Trained: a difference of 18.3 percentage points. The paper does not report confidence intervals or a statistical-significance test for this downstream table. These 15-trial, task-specific results support evaluation within the reported setup; they do not establish universal performance or generalization across other hands, tactile systems, objects, controllers, or operating conditions.

Engineering implications

The scale of HT-Bench makes data organization as important as model architecture. Millions of adjacent frames can still be highly correlated, so teams need task-level and trajectory-level splits that prevent similar contact sequences from appearing in both training and evaluation. Sensor identity, hand geometry, calibration, and task boundaries must stay visible in the metadata.

The benchmark also reinforces why full-hand tactile sensing is different from a fingertip demonstration. Palm and finger contacts create a distributed observation whose geometry depends on the hand pose, object pose, and contact sequence. The tactile dataset directory explains how to preserve those collection units, while the tactile sensor benchmark guide separates representation evidence from hardware-selection evidence.

What this means for humanoid and dexterous hands

Egocentric vision and full-hand touch are complementary. Vision supplies object and scene context, while tactile sensing records contact hidden by the hand itself. A representation benchmark should test both alignment and independence: touch should agree with visible context when appropriate but still carry useful information when vision is occluded.

For teams evaluating humanoid hands, HT-Bench provides a research checklist rather than a procurement score. Compare it with the humanoid robot skin guide to map benchmark tasks onto coverage, routing, synchronization, replacement, and robot-control constraints.

What this does not prove yet

HT-Bench remains an arXiv preprint centered on one reported egocentric/full-hand tactile sensing pipeline. The authors explicitly list fingertip optical tactile sensors, force/torque sensors, skin-like taxel arrays, and non-hand embodiments as sensing or embodiment categories the benchmark does not yet cover. It therefore does not establish a universal tactile representation for every hand, skin, sensor, or manipulation task. Its reported improvements are tied to the paper's data, baselines, model, task definitions, and evaluation protocol. Independent reproduction and transfer to other systems remain necessary.

The scale figures also should not be interpreted as independent sample counts. RGB and tactile frames from the same task trajectory share temporal and physical context. Evaluation quality therefore depends on how tasks, trajectories, objects, hands, and sensors are separated across splits.

Version 2 says the authors will release the data, evaluation protocols, pretrained weights, and training/testing scripts. As of the August 22 review, the arXiv record did not provide a dedicated downloadable package or official repository for those artifacts. The paper's CC BY 4.0 article license should not be read as a verified license for unreleased dataset or model files.

Evaluation checklist

  • Verify which robot hands, tactile layouts, and camera views produced the data.
  • Preserve task and trajectory boundaries when creating training and test splits.
  • Report contact-geometry, cross-modal, temporal, and downstream-control results separately.
  • Test unseen tasks, objects, sensor units, and embodiments when making transfer claims.
  • Compare learned representations with raw-signal, vision-only, and no-tactile baselines.
  • Treat source-reported metrics as benchmark evidence, not deployment certification.

Source

arXiv v2 abstract: HT-Bench: Benchmarking and Learning Dexterous Full-Hand Tactile Representations with Egocentric Vision

arXiv v2 HTML, including Tables 2 and 4

Check the data behind this research

Compare the reported collection with the publicly listed files, dataset license, split documentation and dated access evidence.

Read the directory’s public availability and reproduction analysis.

Continue the topic

Tactile AIUniVTAC separates tactile simulation, representation learning, and policy evaluationTactile AIVision-based tactile intelligence connects sensor optics to robot actionTactile AIFeelWorld predicts contact, tactile force states, and slip for robot planning