HumanoidToolBench separates tool choice from task completion
HumanoidToolBench tracks contact, correct selection, lifting and final success across stationary and mobile Unitree G1 tool-use tasks instead of collapsing them into one score.

Researchers from Seoul National University, the University of Massachusetts Amherst and Google Research released HumanoidToolBench on October 1, 2026. The benchmark asks a Unitree G1 to infer which tool can accomplish a goal, grasp it and use it, sometimes while moving along a table. Its 18 tasks, 55 tool assets and 3,094-trajectory ToolBook dataset make the hand-to-object chain measurable instead of treating “tool use” as a single opaque success bit. Paper and version record.
Key takeaways
- Three scenarios—BallMove, BallRetrieve and IceBreak—combine with three execution levels and standard or decoy tool sets to form 18 conditions.
- ToolBook contains 3,003 simulation trajectories and 91 curated real-robot trajectories collected by eight experienced teleoperators.
- The released evaluation toolkit is MIT-licensed, while ToolBook is CC BY-NC 4.0 and third-party assets retain their own terms.
An evaluation funnel for physical tool use
Each scene contains one suitable tool, two irrelevant objects and, in decoy mode, a misleading tool that differs in task-relevant length, shape, mass or compliance. Instructions describe the goal but do not name the tool. Level L0 ends when the correct tool is lifted 8 cm. L1 adds stationary execution. L2 moves the task target so the robot must locomote while carrying the tool.
The environment records five events: any candidate contact, correct first contact, correct-tool lift, tool-target contact and final success. That funnel separates visual selection from grasping and contact execution. “Correct first contact” is still a proxy—it does not prove semantic reasoning—but it shows where a policy begins to diverge from the intended sequence.
ToolBook is built from 1,959 L1/L2 simulation attempts that yielded 1,200 successful demonstrations. The authors cut 1,803 L0 prefixes from attempts in which the correct tool was lifted, including 614 attempts that later failed. Adding 91 real L1 demonstrations gives exactly 3,094 trajectories. Dataset construction and protocol.
What the scores reveal
Simulation evaluation uses one trained policy instance and 100 episodes per task. FastWAM's standard-mode L1 success is 76%, 57% and 86% across the three scenarios, but its mobile L2 results fall to 9%, 20% and 69%. In decoy mode, the corresponding L1 scores are 75%, 45% and 82%; L2 scores are 8%, 11% and 61%. The benchmark changes both target layout and locomotion demand at L2, so the drop cannot be attributed to walking alone.
Averaged across scenarios and L1/L2, FastWAM contacts some candidate in 99.8% of standard episodes and makes the correct tool its first contact in 87.7%, yet completes only 52.8%. Under decoys those values are 100%, 74.7% and 47.0%. Contact is easy; preserving the right grasp and applying the tool is the larger gap.
Real evaluation covers L1 BallMove and BallRetrieve, each in two modes, with ten trials per cell. GR00T N1.7 succeeds in 18 of 40 trials; FastWAM also succeeds in 18 of 40; ACT succeeds in four. These are small platform-specific trials, not estimates of general humanoid tool competence.
RoboSkin analysis
The strongest contribution is the intermediate event record. It lets researchers connect robot-hand contact to final behavior: a policy may visually select the right object but fail at lift, maintain contact yet miss the target, or complete a stationary task and fail once whole-body motion is introduced.
Focused probes also warn against over-interpreting success. GR00T N1.7's correct first-contact rate falls from 85% to 74% on unseen BallMove tools and from 59% to 36% on BallRetrieve. In one fixed IceBreak scene, an unrelated instruction still triggers task execution in 84 of 100 rollouts, compared with 77 under the intended instruction. That is evidence of scene-policy shortcuts, not language-grounded tool reasoning.
Availability and replication boundaries
The official repository was live when checked. It includes the evaluation environment, recording validator, ACT and Diffusion Policy checkpoints and documented 100-episode protocol. The maintainers estimate about two hours and 1 GB of outputs for one condition on one GPU, roughly a day and a half and 15 GB for all 18 conditions sequentially. Training pipelines are not included.
The paper is an arXiv v1 preprint. RoboSkin.ai inspected the published paper, project and repository but did not run the GPU benchmark or physical robot tests. Tool assets, third-party models and controller components carry separate terms; users should not infer that the dataset's non-commercial license covers every dependency.


