Research proposalVersion 0.2Not peer reviewed

Interaction as the Interface

A Belief-Space Framework for Generalist Embodied Intelligence

Prepared for RoboSkin.ai · AI-assisted research draft ·

The research question

When is another interaction worth it?

A robot may need to touch, move, or test an object before it can choose a useful action. This proposal asks whether explicitly representing uncertain physical states and predicted action outcomes can improve those decisions.

The manuscript describes an interface connecting observation history, feasible commands, predicted consequences, and feedback. An analytical example and a reproducible synthetic benchmark examine when additional sensing helps, becomes redundant, or causes harm through an incorrect sensor model.

  1. 01

    Observe

  2. 02

    Update belief

  3. 03

    Predict outcomes

  4. 04

    Probe when useful

  5. 05

    Act and replan

This loop describes the proposed architecture. The released executable evaluates only the binary decision example.

Current evidence

Analytical example and synthetic benchmark completed. Learning architecture proposed.

Derived and checked

Analytical example

The binary decision rule was checked against explicit outcome enumeration. It is a standard Bayesian information-value calculation.

Executed and reproduced

Synthetic benchmark

Parameters, seeds, code, and numerical outputs are available. A separate local rerun reproduced all five numerical CSV files.

Proposed

Learning architecture

The broader belief, prediction, and control interface has not been implemented or trained as a complete model.

Not evaluated

Robot performance

Physical-robot performance, cross-embodiment transfer, and emergent capabilities have not been demonstrated.

The benchmark checks a standard decision rule under a known observation model. Establishing a distinct algorithmic or theoretical contribution is part of the continuing research.

A small, inspectable decision example

In the specified binary problem, a useful probe must improve the final decision enough to repay its cost. If the best immediate success probability is M, the probe accuracy is p, and its cost is c, the optimal net utility is:

V* = max(M, p − c)

This formula assumes two final actions with reward 1 for a correct choice and 0 otherwise, a known symmetric binary probe with accuracy p ≥ 0.5, one state-preserving probe, and nonnegative cost c measured in the same reward units. It is an illustrative Bayesian calculation.

Measured net utility in the synthetic benchmark
Visual accuracyDirect actionSelective probeDifference
0.500.5003500.7991100.298760
0.750.7496550.7991100.049455
0.950.9503950.9503950.000000

Probe accuracy 0.90; cost 0.10. Twenty seeds reuse 200,000 base random scenarios across three conditions, giving 600,000 episode-condition evaluations. These are synthetic calculations, not physical-robot trials. Within each condition the selective policy either always probes or always skips.

The manuscript and download package include exact expectations, per-seed results, and paired intervals that measure Monte Carlo sampling error only.

The next research question

With limited calibration data, how reliably can a system determine whether a probe will improve the task outcome? The next study will need a precise method and fair comparisons with existing decision-focused experimental design and robust control.

Relevant precedents include vision-and-touch grasp adjustment, GoBOED, and distributionally robust POMDPs. The proposal does not establish a new general control principle.

Version history

  1. Version 0.2 27 September 2026 · Current download

    Clarifies proposal status, corrects belief-update and value-approximation details, adds direct precedents, explains shared random scenarios, and records a local reproducibility check of the numerical outputs.

  2. Version 0.1 27 September 2026 · Initial internal draft

    Initial proposal, analytical decision example, and synthetic benchmark. Superseded by version 0.2.