GIST humanoid visual-tactile-action dataset maps 101.9K soft-object samples
The GIST preprint reports 101.9K visual-tactile-action samples for towel and sponge manipulation, with dense hand touch, two camera views, and explicit access limits.

Updated technical brief - August 2026
The GIST v2 preprint reports 101.9K visual-tactile-action samples collected by three experimenters across four towel and sponge pressure conditions. Its two Inspire RH56-DFX hands provide 2,124 tactile sensing units in total. The source is useful as a documented soft-object collection and ACT baseline study, but it is not currently a verified downloadable dataset or cross-humanoid benchmark.
Source findings
The paper pairs arm and finger joint positions, egocentric and third-person vision, dense hand pressure signals, and robot actions. The authors normalize the hand-tactile readings from the raw 0-4095 range to 0-1 before analysis. The paper describes a humanoid teleoperation setup but does not identify the humanoid robot model, so no specific platform should be inferred from the collection workflow.
| Collection layer | What arXiv v2 reports | Evidence boundary |
|---|---|---|
| Scale | 101.9K samples | The paper uses both “frames of motion data” and “samples”; it does not provide a public file manifest. |
| Operators | Three experimenters followed the collection protocol | This is limited operator diversity, not population-scale coverage. |
| Robot hands | Two Inspire RH56-DFX dexterous hands | The humanoid robot model is not named. |
| Hand touch | 1,062 tactile sensing units per hand, 2,124 total, distributed over fingers and palms | Sensor calibration, unit-to-unit drift, and replacement procedures are not reported. |
| Egocentric view | Head-mounted camera at 848 × 480 | The paper does not state the head-camera model or frame rate. |
| Third-person view | Intel RealSense D435 positioned about 1 m to the robot's left | One fixed external viewpoint does not establish multi-environment coverage. |
| External pressure | Piezoresistive tactile carpet producing real-time pressure heatmaps | The carpet supports pressure feedback during collection and should not be confused with the 2,124 hand tactile units. |
Four soft-object pressure conditions
The main collection covers two deformable objects under strong and weak handling instructions. Each condition contains approximately 77-80 episodes, and each episode lasts 20-30 seconds. The tactile-carpet heatmap gave the experimenter real-time pressure feedback, but the paper does not define a calibrated numerical threshold separating “strong” from “weak.”
| Condition | Object | Pressure instruction | Reported episode coverage |
|---|---|---|---|
| Towel Strong | Towel | Strong | Approximately 77-80 episodes, 20-30 seconds each |
| Towel Weak | Towel | Weak | Approximately 77-80 episodes, 20-30 seconds each |
| Sponge Strong | Sponge | Strong | Approximately 77-80 episodes, 20-30 seconds each |
| Sponge Weak | Sponge | Weak | Approximately 77-80 episodes, 20-30 seconds each |
The paper also describes rigid-object observations for comparison, but its defined four-task dataset and policy experiments focus on towel and sponge manipulation. That distinction matters: rigid-object comparison figures do not expand the reported task benchmark beyond the two soft-object categories.
ACT dense versus sparse baseline
The authors convert the full 2,124-unit hand signal into a 42-location sparse representation, an approximately 98% reduction, and train Action Chunking Transformer baselines called ACT-Dense and ACT-Sparse. Both use the same four conditions. The paper reports an 80/20 train/evaluation split, 100K training steps, evaluation every 20K steps, three seeds, and Mean Absolute Error over predicted versus ground-truth actions.
| Baseline | Tactile input | What the paper reports | What not to infer |
|---|---|---|---|
| ACT-Dense | 2,124 hand tactile units represented as spatial maps | Training loss declined, while dense signals exposed high dimensionality and noise that made optimization harder. | More taxels did not automatically produce a large downstream policy gain. |
| ACT-Sparse | 42 selected locations | Training and test curves were broadly similar to ACT-Dense, with a relatively small test-loss gap overall. | Sparse touch was not established as universally sufficient for contact-rich manipulation. |
The paper's t-SNE analysis separates the four pressure conditions more clearly with dense than sparse signals. That is representation evidence, not proof of a correspondingly large improvement in real-world task success. The real-world section shows representative successful trials but does not publish a success-rate table for comparing ACT-Dense and ACT-Sparse.
RoboSkin analysis
This work is most valuable as a concrete data-contract example for humanoid touch: hand coverage, two visual viewpoints, proprioception, action, operator protocol, and pressure context all need to be interpreted together. It also exposes an important tactile AI problem: collecting dense touch is easier than learning a robust control advantage from it.
The tactile robotics dataset directory places this collection alongside datasets with different embodiments and availability states. The humanoid robot skin guide maps its hand coverage into the wider body-level tactile stack. The tactile AI pillar explains why sensing density, representation learning, and closed-loop policy evidence must be evaluated separately.
Engineering implications
A usable visual-tactile-action dataset needs more than a large sample count. Teams need exact timestamps, stream rates, calibration records, trajectory boundaries, sensor-layout metadata, train/test split rules, and robot-action semantics. These details determine whether a model learns contact dynamics or correlations tied to one operator, sensor unit, view, or episode.
The dense-versus-sparse comparison also argues for measuring both representation quality and downstream control. Clearer clustering can show that dense signals preserve pressure-condition structure, while policy loss reveals whether the learning system can exploit it. Those are different questions and should remain separate.
Availability and licensing
As of 2026-08-22, no official dataset download, code repository, dedicated project page, or dataset license could be verified from the paper or its arXiv record. The CC BY 4.0 notice on arXiv applies to the article. It does not by itself grant a license to unpublished dataset files, code, images, or other research artifacts.
Evaluation checklist
- Preserve episode-level splits instead of randomly mixing adjacent samples across training and evaluation.
- Report camera and tactile sampling rates, timestamps, clock alignment, and dropped-frame handling.
- Publish the robot model, action schema, hand calibration, tactile layout, and replacement procedure.
- Define strong and weak pressure conditions with repeatable physical measurements.
- Compare dense and sparse touch using downstream success, failure, and recovery metrics in addition to MAE.
- Test new objects, operators, hands, viewpoints, and humanoid embodiments before making transfer claims.
- Verify a dataset-specific download URL and license before describing the collection as open data.
What this does not prove yet
This is one arXiv preprint on two soft objects, four pressure conditions, three experimenters, one unnamed humanoid embodiment, and one dual-hand tactile layout. It does not establish transfer to other humanoids, robot hands, sensor technologies, objects, environments, or pressure definitions. The 101.9K samples are temporally related observations, not 101.9K independent tasks or trajectories.
The ACT study also does not show that dense touch always outperforms a 42-location sparse representation. The authors report relatively small test-loss differences and identify high dimensionality, noise, and optimization as open challenges. Independent reproduction is still needed.
Primary source and access check
arXiv v2: A Humanoid Visual-Tactile-Action Dataset for Contact-Rich Manipulation

