<- Back to research

GIST humanoid visual-tactile-action dataset maps 101.9K soft-object samples

The GIST preprint reports 101.9K visual-tactile-action samples for towel and sponge manipulation, with dense hand touch, two camera views, and explicit access limits.

GISThumanoid visual-tactile-action datasetInspire RH56-DFXsoft-object manipulationdense tactile sensing
Illustration for GIST humanoid visual-tactile-action dataset maps 101.9K soft-object samples

Updated technical brief - August 2026

The GIST v2 preprint reports 101.9K visual-tactile-action samples collected by three experimenters across four towel and sponge pressure conditions. Its two Inspire RH56-DFX hands provide 2,124 tactile sensing units in total. The source is useful as a documented soft-object collection and ACT baseline study, but it is not currently a verified downloadable dataset or cross-humanoid benchmark.

Source findings

The paper pairs arm and finger joint positions, egocentric and third-person vision, dense hand pressure signals, and robot actions. The authors normalize the hand-tactile readings from the raw 0-4095 range to 0-1 before analysis. The paper describes a humanoid teleoperation setup but does not identify the humanoid robot model, so no specific platform should be inferred from the collection workflow.

Collection layerWhat arXiv v2 reportsEvidence boundary
Scale101.9K samplesThe paper uses both “frames of motion data” and “samples”; it does not provide a public file manifest.
OperatorsThree experimenters followed the collection protocolThis is limited operator diversity, not population-scale coverage.
Robot handsTwo Inspire RH56-DFX dexterous handsThe humanoid robot model is not named.
Hand touch1,062 tactile sensing units per hand, 2,124 total, distributed over fingers and palmsSensor calibration, unit-to-unit drift, and replacement procedures are not reported.
Egocentric viewHead-mounted camera at 848 × 480The paper does not state the head-camera model or frame rate.
Third-person viewIntel RealSense D435 positioned about 1 m to the robot's leftOne fixed external viewpoint does not establish multi-environment coverage.
External pressurePiezoresistive tactile carpet producing real-time pressure heatmapsThe carpet supports pressure feedback during collection and should not be confused with the 2,124 hand tactile units.

Four soft-object pressure conditions

The main collection covers two deformable objects under strong and weak handling instructions. Each condition contains approximately 77-80 episodes, and each episode lasts 20-30 seconds. The tactile-carpet heatmap gave the experimenter real-time pressure feedback, but the paper does not define a calibrated numerical threshold separating “strong” from “weak.”

ConditionObjectPressure instructionReported episode coverage
Towel StrongTowelStrongApproximately 77-80 episodes, 20-30 seconds each
Towel WeakTowelWeakApproximately 77-80 episodes, 20-30 seconds each
Sponge StrongSpongeStrongApproximately 77-80 episodes, 20-30 seconds each
Sponge WeakSpongeWeakApproximately 77-80 episodes, 20-30 seconds each

The paper also describes rigid-object observations for comparison, but its defined four-task dataset and policy experiments focus on towel and sponge manipulation. That distinction matters: rigid-object comparison figures do not expand the reported task benchmark beyond the two soft-object categories.

ACT dense versus sparse baseline

The authors convert the full 2,124-unit hand signal into a 42-location sparse representation, an approximately 98% reduction, and train Action Chunking Transformer baselines called ACT-Dense and ACT-Sparse. Both use the same four conditions. The paper reports an 80/20 train/evaluation split, 100K training steps, evaluation every 20K steps, three seeds, and Mean Absolute Error over predicted versus ground-truth actions.

BaselineTactile inputWhat the paper reportsWhat not to infer
ACT-Dense2,124 hand tactile units represented as spatial mapsTraining loss declined, while dense signals exposed high dimensionality and noise that made optimization harder.More taxels did not automatically produce a large downstream policy gain.
ACT-Sparse42 selected locationsTraining and test curves were broadly similar to ACT-Dense, with a relatively small test-loss gap overall.Sparse touch was not established as universally sufficient for contact-rich manipulation.

The paper's t-SNE analysis separates the four pressure conditions more clearly with dense than sparse signals. That is representation evidence, not proof of a correspondingly large improvement in real-world task success. The real-world section shows representative successful trials but does not publish a success-rate table for comparing ACT-Dense and ACT-Sparse.

RoboSkin analysis

This work is most valuable as a concrete data-contract example for humanoid touch: hand coverage, two visual viewpoints, proprioception, action, operator protocol, and pressure context all need to be interpreted together. It also exposes an important tactile AI problem: collecting dense touch is easier than learning a robust control advantage from it.

The tactile robotics dataset directory places this collection alongside datasets with different embodiments and availability states. The humanoid robot skin guide maps its hand coverage into the wider body-level tactile stack. The tactile AI pillar explains why sensing density, representation learning, and closed-loop policy evidence must be evaluated separately.

Engineering implications

A usable visual-tactile-action dataset needs more than a large sample count. Teams need exact timestamps, stream rates, calibration records, trajectory boundaries, sensor-layout metadata, train/test split rules, and robot-action semantics. These details determine whether a model learns contact dynamics or correlations tied to one operator, sensor unit, view, or episode.

The dense-versus-sparse comparison also argues for measuring both representation quality and downstream control. Clearer clustering can show that dense signals preserve pressure-condition structure, while policy loss reveals whether the learning system can exploit it. Those are different questions and should remain separate.

Availability and licensing

As of 2026-08-22, no official dataset download, code repository, dedicated project page, or dataset license could be verified from the paper or its arXiv record. The CC BY 4.0 notice on arXiv applies to the article. It does not by itself grant a license to unpublished dataset files, code, images, or other research artifacts.

Evaluation checklist

  • Preserve episode-level splits instead of randomly mixing adjacent samples across training and evaluation.
  • Report camera and tactile sampling rates, timestamps, clock alignment, and dropped-frame handling.
  • Publish the robot model, action schema, hand calibration, tactile layout, and replacement procedure.
  • Define strong and weak pressure conditions with repeatable physical measurements.
  • Compare dense and sparse touch using downstream success, failure, and recovery metrics in addition to MAE.
  • Test new objects, operators, hands, viewpoints, and humanoid embodiments before making transfer claims.
  • Verify a dataset-specific download URL and license before describing the collection as open data.

What this does not prove yet

This is one arXiv preprint on two soft objects, four pressure conditions, three experimenters, one unnamed humanoid embodiment, and one dual-hand tactile layout. It does not establish transfer to other humanoids, robot hands, sensor technologies, objects, environments, or pressure definitions. The 101.9K samples are temporally related observations, not 101.9K independent tasks or trajectories.

The ACT study also does not show that dense touch always outperforms a 42-location sparse representation. The authors report relatively small test-loss differences and identify high dimensionality, noise, and optimization as open challenges. Independent reproduction is still needed.

Primary source and access check

arXiv v2: A Humanoid Visual-Tactile-Action Dataset for Contact-Rich Manipulation

Continue the topic

Tactile DataFreeTacMan robot-free visuo-tactile data collection for tactile AIRobot learningSlipSense: from pressure and vibration to a timely regraspRobot learningTouch2Trace: what makes a tactile cable-tracing policy work