InterEvolve searches reward programs instead of retraining a humanoid
InterEvolve uses an LLM, numerical tuning and parallel simulation to revise staged reward programs while keeping the Unitree G1 control policy fixed.

University of Illinois Urbana-Champaign researchers released InterEvolve on October 1, 2026. The system adapts a frozen humanoid controller to new contact-rich tasks by changing the reward program that queries it, rather than updating policy weights. An LLM edits program stages, CMA-ES tunes numerical constants, and parallel simulation verifies proposals. The reported macro-average rises from 34.6% for a tuned initial agent program to 86.5% after evolution. Paper and version record.
Key takeaways
- An object-aware forward-backward behavioral foundation model turns rewards into latent prompts for one fixed Unitree G1 controller.
- Full InterEvolve reports 86.5% simulation success, 2.1 GPU-hours and 0.23 million LLM tokens across eight task families under the paper's protocol.
- Real-robot demonstrations use an RGB-D camera on the robot, but perception, state estimation and policy inference run on an off-board GPU workstation.
Evolve the interface, not the weights
The motor model starts with a frozen body prior and adds object-aware residual networks trained on retargeted OMOMO and GRAB interactions. The training collection contains 3,866 clips; 50 clips for each of four box-like objects are held out for tracking evaluation. At test time, a reward is projected against a state bank into a latent prompt that elicits behavior from the fixed policy.
A reward program can contain multiple stages, completion conditions, weights and thresholds. DeepSeek-V4-Flash revises the program structure from simulator feedback, while CMA-ES calibrates constants. Each valid proposal receives 192 tuning rollouts and another 192 confirmation rollouts. Verified programs become a text skill library that later tasks can retrieve. Method, search budget and evaluation details.
This is adaptation through task specification. It can expose behavior already present in the controller, but it does not manufacture a missing motor primitive. The same distinction appears in contact-rich Physical AI: better objectives can coordinate pushing, lifting and carrying only if the underlying control distribution contains usable versions of those motions.
What improved—and what did not
Across eight large-box task families, the initial agent program with numerical tuning reaches 34.6% success. The full method reaches 86.5%. Removing CMA-ES reduces that result to 51.6%; forcing one stage reduces it to 44.7%; removing targeted edits, multi-scenario testing or scene context produces 68.9%, 68.4% and 78.7%.
The progression is not monotonically successful. Full success moves from 32.2% initially to 75.0%, 76.8%, 73.6% and finally 86.5% over four rounds. The smoother “earned criteria” metric rises every round, but a change that fixes one criterion can temporarily break another when all conditions must hold in the same rollout.
Skill-library reuse is tested on three composite tasks with ten episodes each. The full library completes 8/10 relocate, 4/10 stack and 5/10 carry-place-kick episodes. With no library those counts are 0/10, 0/10 and 1/10. These cells are small and task-specific, but they show that verified task descriptions can be reusable artifacts.
RoboSkin analysis
The ablations make stage structure and calibration inseparable. A high-level program such as approach, grasp, carry and release is not sufficient if its thresholds do not match the robot and scene. Conversely, numerical tuning cannot repair a single-stage objective that confuses contact acquisition with task completion. This is directly relevant to humanoid manipulation, where each phase changes which contacts are desirable.
The 86.5% figure is a simulation macro-average, not a hardware success rate. The project page shows autonomous G1 executions, including repeated kicks, but one kick sequence uses motion-capture box pose. Other demonstrations use robot-mounted RGB-D input with off-board computation. The paper does not publish a broad table of repeated physical success counts comparable to the simulation study.
Limitations and availability
InterEvolve spends inference-time compute on LLM calls, CMA-ES and hundreds of simulator rollouts per candidate. It does not replan reward code in real time during hardware execution. Results are bounded by the pretrained motor repertoire, reward feature bank, available scene measurements and simulator fidelity. Transfer to different objects can require manual program adaptation; unchanged large-box programs show large drops on several objects.
This is an arXiv v1 preprint, and RoboSkin.ai did not run the simulator or robot. The official project publishes explanatory material and videos. No official source-code repository, trained controller, dataset download or software license was visible when checked. The manuscript is distributed under arXiv's perpetual non-exclusive license.


