Body-grounded replanning lets joint load change the robot strategy
Microsoft Research Asia, Waseda, Chiba and NII researchers move proprioceptive evidence into high-level replanning. Hardware success reaches 96.3%, while LLM pause time is excluded.

Researchers at Microsoft Research Asia–Tokyo, Waseda University, Chiba University and Japan's National Institute of Informatics released Body-Grounded Replanning on September 24, 2026. The framework turns joint position, actuator effort, contact and recent execution history into high-level evidence: when the current motion becomes physically unsuitable, GPT-4o-mini selects another strategy while the task objective and low-level controller remain unchanged. Paper and version record.
Key takeaways
- In ten real-robot reaching conditions, body-grounded replanning and an environment-only search baseline both reach 96.3% success; the main advantage is faster restricted completion, 24.57 seconds versus 28.18 seconds.
- A rule-based controller using the same body-state events reaches 66.3%, showing that detecting strain or contact is not equivalent to choosing the right alternative.
- The three contact-rich tasks—button pressing, box-lid opening and lint-roller extraction—are qualitative demonstrations without repeated success rates.
What changed
Most manipulation planners ask whether a trajectory is geometrically feasible. This work adds a second question: is the strategy physically suitable for the robot's current body state? A path may still reach the target while producing high joint effort, approaching a joint limit or encountering contact earlier than expected.
The monitor records per-joint position, velocity and effort plus binary contact at 20 Hz. Current measurements and a learned short-horizon prediction trigger events such as high effort, limited mobility, no progress or early contact. GPT-4o-mini, run at temperature zero, receives the current strategy, joint-level state, recent execution statistics, event summary, history and a manually defined menu of alternative strategies. Its structured choice changes the high-level motion while leaving the low-level controller in place.
Results under the reported conditions
The controlled simulation uses a Franka Panda and generates 100 paired episodes for each of ten load or asymmetric-mobility conditions. The body-grounded method maintains 90% to 99% success and records the lowest mean joint-torque norm in every condition. Across the seven asymmetric-mobility conditions, overall success is 96.7%, compared with 88.5% for the rule-based body-state baseline. Full results.
The hardware study uses a CRANE-X7 with Dynamixel actuators. Environment-only search, rules and body-grounded replanning each receive eight trials per condition under a 45-second execution budget. The body-grounded and environment-only methods both average 96.3% success; rules average 66.3%. Body grounding has the lowest restricted completion time in nine of ten conditions and the lowest overall mean at 24.57 seconds. Its average actuator-effort norm is also lowest—2.18 versus 2.28 and 2.33—but it wins that metric in only five of ten individual conditions.
Restricted completion time assigns the full 45 seconds to failures and excludes LLM inference pauses. That makes the metric useful for combining reliability and active robot time, but it is not end-to-end wall-clock latency. The paper does not publish a separate distribution for network or model response time.
The authors also show one execution each for obstacle-avoiding button pressing, box-lid opening and lint-roller extraction. Contact or impeded motion triggers a strategy change and task completion in the shown sequences. These examples demonstrate interface reuse, not a quantified manipulation benchmark.
What this means for robotics
RoboSkin analysis: body-grounded replanning moves proprioceptive and contact signals above the servo layer. Instead of only correcting torque locally, the robot can decide that the entire approach direction, posture or action order is wrong for its current physical condition.
The comparison with environment-only search is equally important. Both succeed at the same rate on hardware, so the evidence supports efficiency and physical suitability more strongly than raw completion. A production system would also need bounded fallback behavior when the language model is slow, unavailable or selects an unsafe strategy.
Limitations and availability
This is an arXiv v1 preprint. The high-level strategy set and event thresholds are manually defined, tasks are short-horizon, and the strongest repeated evaluation is controlled reaching rather than the three contact-rich demonstrations. Hardware effort uses raw Dynamixel readings, not calibrated joint torque, so values are only compared within that platform.
The paper links an author-provided qualitative video, but no public code repository, prompt package, logs, dataset or software license was verified. GPT-4o-mini is an external hosted model, and the evaluation excludes its inference pauses from active execution time. RoboSkin.ai has not reproduced the system.


