<- Back to news

Uni-VLaT gives humanoid VLA policies whole-body touch

Uni-VLaT adapts pretrained humanoid VLA policies with textile electronic skin across eight body regions and predicts future touch, body state and visual features during training.

Preprint · arXiv v1 · five real-robot tasks · project videos available · code not releasedSource date: Read the primary source ↗
whole-body tactile sensinghumanoid VLA adaptationelectronic skinloco-manipulation
Diagram of eight tactile regions on a humanoid feeding a VLA policy and three future-representation prediction heads.
Original RoboSkin.ai schematic of Uni-VLaT whole-body tactile adaptation. It is explanatory artwork, not an experimental image.

Researchers from Tsinghua University, Beihang University, the Communication University of China and the University of Hong Kong released Uni-VLaT on September 28, 2026. The method adapts pretrained vision-language-action policies to a Unitree G1 humanoid with distributed textile electronic skin, then uses touch as an anchor for predicting future tactile, proprioceptive and visual representations. In five real-robot tasks, the authors report a 75% average success rate. Paper and version record.

Key takeaways

  • Uni-VLaT reaches 75% average success over five tasks, versus 32% without tactile input, 68% with touch but no prediction and 69% with tactile-only prediction.
  • The main Isaac-GR00T comparison uses 20 physical rollouts per configuration and about 50 demonstrations per task; the separate pi-0.5 comparison uses 10 rollouts per configuration.
  • The tactile system covers eight torso, shoulder, back and arm regions. Its values are pressure-related ADC readings in arbitrary units, not calibrated force measurements.

What changed

Appending touch to a VLA policy does not guarantee that the model connects contact with body motion and scene change. Uni-VLaT encodes a four-frame tactile history into eight learned tokens, mixes those tokens with proprioceptive and action tokens inside the policy, and adds three training-only heads. Those heads predict four future tactile, proprioceptive and visual representations.

The deployed system keeps the tactile pathway and action generator but discards the predictive heads. A pretrained SONIC controller converts each generated 64-dimensional motion token into coordinated whole-body commands. The vision-language model stays frozen while the action expert, state encoder and tactile modules are fine-tuned. Method details.

The hardware is a Unitree G1 with stereo RGB, proprioception and custom JQ Industries textile electronic skin. Sensors cover the chest, central back, both shoulders, both upper-back regions and both arms. Circumferential sleeve readings are pooled to reduce sensitivity to sleeve rotation, which also reduces the spatial detail available around the arm.

Results under the reported conditions

With Isaac-GR00T as the policy backbone, the no-touch baseline averages 32% across Back-Tap Walking, Table Sweeping, Basket Loading, Human-Robot Hugging and Composed Cleanup. Adding tactile input without prediction raises the average to 68%. Tactile-only future prediction reaches 69%, and the complete three-modality objective reaches 75%.

The task breakdown matters. Back-Tap Walking rises from 0% without touch to 85% with Uni-VLaT because contact itself is the instruction to walk or stop. Table Sweeping rises from 45% to 75%, while Basket Loading rises from 30% to 80%. Human-Robot Hugging reaches 80%, and the three-stage Composed Cleanup task reaches 55%. Each headline percentage in the main table corresponds to 20 real-robot rollouts. Evaluation table.

The cross-backbone study reports the same directional gains with pi-0.5: 0% to 90% on Back-Tap Walking, 30% to 60% on Table Sweeping and 30% to 60% on Basket Loading. Those configurations use ten rollouts, so their percentage resolution is ten points. A from-scratch diffusion-policy baseline is not assigned a success rate because three checkpoints failed the authors' SONIC pre-execution safety screen.

What this means for robotics

RoboSkin analysis: whole-body tactile sensing changes what a humanoid can observe, especially when contact is broad, occluded or outside the hands. The important result is not merely the 43-point gap over the no-touch baseline. The smaller 7-point gap over touch without prediction suggests that most of the gain comes from making contact directly observable, while multimodal future supervision adds a further task-dependent benefit.

For teams building electronic skin systems, the study also illustrates a practical representation choice: divide the body into stable regions, preserve temporal change, and avoid treating every taxel as an interchangeable flat vector. The trade-off is that pooling and arbitrary-unit readings do not yield calibrated forces or fine contact mechanics.

Limitations and availability

Uni-VLaT is an arXiv v1 preprint, and RoboSkin.ai has not reproduced it. Evaluation uses one humanoid, five tasks and about 50 demonstrations per task. The paper reports rollout counts but not confidence intervals. Results mix a discrete touch-triggered behavior with sustained manipulation, so the five-task average should not be read as one uniform capability measure.

The official project page provides task videos but still labels code as “coming soon.” It also retained stale “paper coming soon” text when checked, even though the arXiv paper was already public. No repository, dataset, checkpoint or software license was verified. The arXiv manuscript license does not grant rights to unreleased implementation assets.

Continue the topic

Tactile robot learningA tactile reflex becomes the teacher for fragile graspingForce-aware VLAVisForce draws force goals into a dexterous robot policyVisuo-tactile dexterityOccluDex fuses 3D geometry and touch when robot hands block the view