<- Back to news

ActiveWAM makes camera motion part of the robot action

ActiveWAM treats where a robot looks as part of the same action sequence as bimanual manipulation, while using transformed video histories only during training.

Preprint · arXiv v1 · 50-task simulation benchmark · 300 physical trials across five methods · official code marked coming soonSource date: Read the primary source ↗
active visionworld-action modelbimanual manipulationcamera control
Diagram showing camera history and task evidence flowing into a world-action model that jointly commands a pan-tilt head and two robot arms.
Original RoboSkin.ai schematic of ActiveWAM's retain-and-acquire loop. It is explanatory artwork, not an experimental image.

Beijing Institute of Technology researchers released ActiveWAM on October 1, 2026, a world-action model that predicts bimanual motion and a two-axis pan-tilt camera trajectory together. The central engineering question is simple: when a robot can move its eyes, can it learn where to look without losing evidence it saw a moment ago? ActiveWAM answers with a finite video history, training-time history transformations and one 16-dimensional arm-and-head action space. Paper and version record.

Key takeaways

  • RoboTwin-AV extends all 50 RoboTwin 2.0 tasks with executable camera control and evaluates 100 episodes per task, seed and condition across three seeds.
  • Under combined appearance and initial-camera-pose shifts, ActiveWAM reports 53.3% success versus 33.3% for Fast-WAM, a 20.0 percentage-point difference.
  • On three physical kitchen sequences, ActiveWAM completes 30 of 60 trials; the system still uses prescribed stage switching and a fixed-focal-length camera.

Retain old evidence, acquire the next view

The policy receives nine 240×320 RGB frames, robot state, camera pose, calibration and acquisition metadata. It produces two six-degree-of-freedom arm increments, two gripper commands and two camera increments. A stay command lets the camera hold a useful view rather than move continuously.

During training, a frozen Wan2.2-TI2V-5B video prior partially transforms recorded histories. Task-weighted feature preservation and temporal correspondence losses constrain the transformation so important cues and visible motion remain. The original action continuation is kept as the target. At deployment, the system uses raw camera history and does not decode a predicted future video. In other words, the video model is a co-training instrument, not a runtime visual simulator. Method and deployment design.

The authors also introduce RoboTwin-AV. Its gaze collector can use simulator-only masks and depth to create active-vision demonstrations, but those privileged inputs are withheld at evaluation. Each of the 50 tasks uses 100 successful training trajectories. Test failures and timeouts remain in the 100-episode evaluation sets.

Results under the reported protocol

ActiveWAM reaches 80.3% in the clean RoboTwin-AV condition, 60.7% under appearance shifts, 65.0% under initial head-pose shifts and 53.3% when both shifts are applied. Fast-WAM scores 76.3%, 41.7%, 56.0% and 33.3%, respectively. Removing history inversion reduces the compound result to 41.7%, an 11.6-point gap; removing the full history reduces it to 34.0%.

The physical setup is an AirbotPlay fixed-base dual-arm platform with a RealSense D455 on a two-DoF head. Three kitchen sequences each receive 20 attempts per method: cucumber slicing, egg frying and ingredient mixing. ActiveWAM completes 7, 11 and 12 full sequences, or 30 of 60. Fast-WAM completes 1, 6 and 7, or 14 of 60. That is a 26.7-point absolute difference, but it does not represent an unattended end-to-end kitchen system: policies are separately trained for segments, transitions are prescribed, and human inter-segment actions are outside the evaluation.

RoboSkin analysis

ActiveWAM is not a tactile system. Its value for robot world models is the treatment of sensing motion as an action with physical consequences. A pan-tilt camera can reacquire an occluded target, but leaving a useful view also creates a memory problem. This retain-and-acquire framing is relevant to visuo-tactile manipulation, where cameras and touch sensors likewise reveal different parts of a contact state over time.

The failure audit is as useful as the headline result. Of 30 failed ActiveWAM physical trials, 12 are labeled manipulation errors, 10 premature terminations, five camera field-of-view failures and three hardware issues. Moving the camera does not remove contact-control or termination problems; it changes which failures remain visible.

Limitations and availability

This is an arXiv v1 preprint, and RoboSkin.ai did not reproduce the experiments. The physical evaluation uses one fixed-base platform, one camera arrangement and three composed kitchen sequences. Inference is reported at about 165 ms on an RTX 4090; action chunking yields roughly 15 Hz effective updates while interpolation maintains 25 Hz commands. The authors report only a short history window and no automatic high-level stage transition.

The official project page was available when checked, but its code link was marked “Coming Soon.” RoboTwin-AV materials are linked from the project, while ActiveWAM training code, checkpoints and a software license were not verified as released. The arXiv manuscript uses arXiv's perpetual non-exclusive distribution license.

Continue the topic

Robot world modelsSkeleWAM compresses manipulation into a sparse 3D skeletonTactile world modelsTacDyn-WAM predicts contact dynamics without generating future touch pixelsContact-aware world modelsInternW0 links asynchronous world prediction with contact-aware pipetting