Robotics datasets: data for robot learning and evaluation

Compare robotics datasets by robot, task, modality, action space, timing, access, and license. Find robot learning, manipulation, teleoperation, VLA, and humanoid data.

Published 2026-08-21 | Updated 2026-08-22 by

Organized robot skin learning library with technical cards, tactile sensor samples, and research screens.
Resource-library visual for public learning routes and technical references.
6
sections
5
questions
8
next routes

Short answer

What you need to know

  1. 1

    A robotics dataset is a structured collection of robot observations, states, actions, task context, outcomes, or environment records used to train, evaluate, or reproduce robotic systems.

  2. 2

    Useful dataset scale is not only the number of frames or trajectories. Robot embodiment, action interface, task diversity, failures, calibration, timing, train-test splits, access, and license determine what can be learned or compared.

  3. 3

    This page covers broad robot learning, manipulation, teleoperation, VLA, and humanoid data. RoboSkin.ai’s /datasets directory remains the specialized hub for tactile and visuo-tactile robotics datasets.

Topic 01

Robotics datasets and tactile datasets serve different scopes

Broad robotics datasets may contain cameras, depth, proprioception, actions, language, rewards, demonstrations, or simulator state without any surface touch. A tactile dataset requires measured tactile data or a clearly defined contact-related modality and should document the sensor and synchronization contract.

RoboSkin.ai therefore keeps two canonical hubs. This page maps the broad robot-data ecosystem; /datasets provides tactile-specific records and filters. Cross-modal resources can appear in both only when each page adds distinct fields and explanation rather than duplicate copy.

HubPrimary scopeMinimum useful fieldsUse it when
Robotics datasetsRobot learning, manipulation, teleoperation, VLA, humanoid, simulation, and evaluation dataRobot, task, observations, actions, collection policy, trajectories, splits, access, licenseComparing broad training or evaluation resources
Tactile robotics datasetsTouch, pressure, force-related, tactile image, slip, contact, and visuo-tactile dataSensor, placement, calibration, rate, synchronization, robot, task, actions, splits, licenseA claim depends on measured touch or contact state

Topic 02

The fields every robot dataset record should expose

A dataset should be understandable before download and reproducible after it. Counts need a unit—episodes, trajectories, frames, steps, hours, tasks, or robots—and those units are not interchangeable. Access and license should be checked against the current repository or dataset card rather than inferred from the paper abstract.

  • Identity: dataset name, institution, version, release date, paper, repository, dataset URL, and maintainers
  • Embodiment: robot, hand or gripper, kinematics, controller, action space, and hardware version
  • Observations: cameras, depth, language, proprioception, force, touch, audio, events, calibration, and units
  • Collection: teleoperation or policy, operator instructions, task definitions, resets, failures, interventions, and quality control
  • Structure: episode unit, counts, timing, synchronization, file format, schema, train-test splits, and checksums
  • Governance: access status, license, usage restrictions, privacy, safety, known gaps, and update history

Topic 03

Source-backed dataset ecosystem map

The resources below illustrate different dataset strategies. The descriptions are based on their primary papers or official project materials; they are not a single leaderboard.

ResourceScopePrimary-source signalBoundary to verify
PRISMContact-rich industrial manipulation across Franka, Realman, and a source-literal LEJU upper-body humanoid setupThe 2026 preprint reports 5,000+ robot trajectories, paired human demonstrations, 45+ hours, and three teleoperation interfacesTactile signals cover only a subset, and the official dataset control remained marked “soon” on 2026-08-22
Open X-EmbodimentAggregated multi-embodiment robot learning dataResearch paper describes cross-institution robot data and generalist-policy experimentsIndividual source datasets, embodiments, action mappings, and licenses differ
DROIDLarge real-world robot manipulation datasetPaper documents diverse manipulation collection across settings and tasksRead the current project and dataset records for exact access, version, schema, and license
LeRobot datasetsOpen tooling and standardized dataset format for robot episodesOfficial documentation and Dataset v3 materials describe storage and loading conventionsQuality and rights remain dataset-specific; format compatibility is not capability evidence
RoboTacDexHumanoid visual-tactile-action data on Unitree G1The 2026 preprint reports 6,000 trajectories, 19 tasks, 23 skills, and 22 objectsThe abstract says open sourcing is forthcoming; verify present access and license before claiming availability
HRDexDBHuman and multi-robot-hand grasp dataThe 2026 preprint reports 1,400 grasps across 100 objects with tactile, visual, and kinematic recordsScale, hand coverage, data quality, splits, and downstream transfer remain protocol-specific

Topic 04

Teleoperation, demonstrations, corrections, and autonomous rollouts

Teleoperation can capture action demonstrations and recovery behavior, but the dataset also contains the operator interface, latency, viewpoint, embodiment constraints, and skill distribution. Autonomous rollouts can add policy-generated successes and failures, while human corrections can target states the current policy handles poorly.

Hugging Face’s LeRobot v0.6 announcement describes rollout tooling and a DAgger-style human-correction workflow. That is an official software capability description, not evidence that every resulting dataset is high quality or that a trained policy will generalize to a new robot.

Topic 05

Training data and benchmark data should not leak into each other

A dataset can support pretraining, fine-tuning, offline evaluation, simulation, or a shared benchmark. The split must match the claim. Random frame splits can leak nearly identical moments from one trajectory into train and test sets; repeated objects or scenes can also inflate an apparent generalization result.

ClaimUseful held-out unitWhat to disclose
New visual conditionsScene, camera, lighting, or backgroundWhich conditions changed and whether geometry or tasks repeated
New objectsObject identity or categorySeen categories, object instances, poses, and physical properties
New tasksTask or skill compositionInstruction templates, subtasks, rewards, and overlap with training demonstrations
New embodimentsRobot, hand, gripper, sensor, or action interfaceRetargeting, calibration, adaptation, and embodiment identifiers
Tactile transferSensor, mounting, material, object, and contact regimeTouch preprocessing, rate, synchronization, drift, and no-touch baseline

Topic 06

How RoboSkin.ai will maintain dataset records

Dataset pages should separate paper claims, repository facts, and editorial normalization. A paper may announce intended release; only the live project or repository can confirm current access. A license field should be copied from a current authoritative source, with version and retrieval date where possible.

Records should retain unknown values, version changes, and source links. This prevents missing metadata from being silently converted into facts and makes the directory more useful for both search engines and AI retrieval systems.

  • One stable entity page per dataset, linked to topics, papers, sensors, robots, tasks, and benchmarks
  • Filterable normalized fields backed by visible primary-source citations
  • Clear labels for available, announced, restricted, archived, or unknown access
  • Change history for schema, URLs, license, checksums, and source status

Source-reviewed robotics dataset directory

Compare general robot-learning datasets

These records cover cross-embodiment and manipulation datasets rather than tactile datasets. Scale, hardware, modalities, formats, access, and licenses remain tied to the reviewed primary sources and version boundaries.

Source review: 2026-08-22 / 3 records

Showing 3 of 3 datasets

DatasetInstitution / yearRobot / sensorModalities / scaleTasks / objectsFormat / licensePrimary links
Open X-Embodiment DatasetPublic dataset inventory, cloud-download instructions, TFDS/RLDS loading examples, and model resources are linked from the official project and repository. This is a versioned aggregation rather than one sensor-homogeneous dataset.Reviewed 2026-08-22Open X-Embodiment Collaboration: 21 institutions and 34 contributing robotics laboratories reported by the primary paper
2023
Robot: 22 robot embodiments across 60 constituent datasets, including single-arm, bimanual, and quadruped platforms
Sensor: Dataset-dependent cameras and sensors; RGB, depth, and point-cloud coverage varies by constituent dataset, with no standardized tactile channel
RGB image; Robot state; Robot action; Task or language annotation when available; Depth or point cloud in subsets
Scale: The current Open X-Embodiment paper reports more than 1 million real-robot trajectories pooled from 60 datasets, covering 22 embodiments and 527 skills. Later model papers use different mixture snapshots, so their larger counts are not substituted for this dataset record.
Cross-embodiment robot manipulation; Language-conditioned robot control; Generalist policy pretraining and adaptation
Objects: Common household and manipulation objects distributed across the constituent datasets; the aggregation does not publish one uniform object taxonomy.
RLDS episodes serialized as TFRecord. Camera counts, action spaces, state fields, task labels, and optional depth or point clouds remain dataset-specific.
License: The official repository licenses software under Apache-2.0 and other repository materials under CC BY 4.0. Users are also instructed to cite the constituent datasets, whose versions and component-level terms should be checked before reuse.
DROID: Distributed Robot Interaction DatasetThe full RLDS dataset and raw data are publicly documented through official Google Cloud paths. The project also links a visualizer, quick-start notebook, hardware documentation, policy code, and updated annotations.Reviewed 2026-08-22DROID Dataset Team: 18 research laboratories represented in the current paper; collection was distributed across North America, Asia, and Europe
2024
Robot: Franka Emika Panda 7-DoF arm with Robotiq 2F-85 gripper on the standardized DROID platform
Sensor: Two external ZED 2 stereo cameras; Wrist-mounted ZED Mini stereo camera; Meta Quest 2 headset and controllers for teleoperation
Stereo RGB video; Depth; Camera calibration; Robot joint and end-effector state; Robot action; Gripper state; Natural-language instruction; Episode and scene metadata
Scale: 76,000 successful demonstration trajectories, about 350 hours, 564 scenes, 86 tasks, 52 buildings, and 50 collectors. The release also includes roughly 16,000 failed trajectories that are not counted in the 76,000 headline.
In-the-wild object manipulation; Language-conditioned imitation learning; Cross-scene and cross-object generalization; Target-domain co-training
Objects: A long-tailed set of everyday objects across household, office, laboratory, and other real-world scenes; no single exhaustive category inventory is claimed here.
The official training release is provided in RLDS format, with a roughly 1.7 TB full set and a 100-trajectory example. Full-HD raw data and optional HDF5 training instructions are documented separately; later language and calibration annotations are published on Hugging Face.
License: CC BY 4.0 for the DROID dataset as stated by the current paper. The official policy-learning code is MIT licensed; that software license is not treated as the dataset license.
BridgeData V2The official project links the raw dataset, a processed TFDS/RLDS release, training code, pretrained checkpoints, and hardware setup instructions.Reviewed 2026-08-22University of California, Berkeley; Stanford University; Carnegie Mellon University; Google DeepMind
2023
Robot: Trossen WidowX-250 6DOF arm
Sensor: Fixed over-the-shoulder RGB-D camera; Two position-randomized RGB cameras; Wrist RGB camera; multi-view and depth coverage varies by trajectory
RGB image; Depth in a subset; Robot state; Robot action; Natural-language task label
Scale: 60,096 trajectories across 24 environments and 13 skills: 50,365 teleoperated demonstrations plus 9,731 scripted pick-and-place rollouts. The paper reports interactions with more than 100 objects.
Pick and place; Pushing and reorientation; Sweeping; Drawer and door interaction; Object stacking; Cloth folding; Granular-media manipulation
Objects: More than 100 objects across toy kitchens, tabletops, sinks, laundry setups, and other manipulation environments.
The raw release contains JPEG, PNG, and pickle files. An official 256 x 256 TensorFlow Datasets version uses RLDS; raw and processed inventories should not be assumed to have identical file counts.
License: CC BY 4.0 for the dataset as stated in the primary paper; the official training repository is MIT licensed.

Common questions

FAQ for this topic

01

What is a robotics dataset?

It is a structured collection of robot observations, states, actions, task context, outcomes, or environment records used for training, evaluation, or reproducibility.

02

What makes a robot learning dataset useful?

It should document the robot, observations, action space, tasks, timing, calibration, collection policy, failures, episode unit, splits, access, license, and known limitations.

03

Is the number of frames enough to compare robotics datasets?

No. Frames, steps, trajectories, episodes, hours, tasks, objects, and robots measure different things. Diversity, quality, synchronization, failures, splits, and rights also determine usefulness.

04

Where are tactile robotics datasets on RoboSkin.ai?

Use /datasets for the specialized tactile and visuo-tactile directory. This /robotics-datasets page covers the broader robot-learning, manipulation, teleoperation, VLA, and humanoid ecosystem.

05

Can an announced dataset be described as open?

Only after a current authoritative source provides access and a license. A paper saying that data will be released is not proof that it is presently downloadable or reusable.