OccluDex fuses 3D geometry and touch when robot hands block the view
OccluDex pretrains a 3D visuo-tactile encoder on human demonstrations, then freezes it for dexterous policies operating under hand-induced occlusion.

Researchers from Shanghai Jiao Tong University, Tongji University and the Chinese University of Hong Kong, Shenzhen released OccluDex on September 30, 2026. The framework combines partial egocentric point clouds with local touch so a dexterous hand can continue estimating object geometry and contact while its own fingers hide the scene. The authors report 17/20 successful physical faucet trials and 14/20 object-reorientation trials after simulation training, without real-world fine-tuning. Paper and version record.
Key takeaways
- A hierarchical masked autoencoder learns global 3D structure at multiple scales, then fuses 19 tactile contact tokens with high-level geometric features.
- Pretraining uses 1,690 human manipulation sequences containing 409,035 synchronized visual-tactile frames from three subjects.
- Simulation averages 100 evaluation episodes per setting and three policy seeds. Physical evaluation uses two unseen objects per task and ten trials per object, giving 40 trials in total.
What changed
An egocentric camera stays aligned with the manipulating body, but the hand itself often blocks the contact region. Completing the hidden shape from vision can be ambiguous, while touch alone is local and says little about the rest of the object. OccluDex gives the two signals different jobs: multiscale point-cloud tokens preserve global geometry, and sensor-specific tactile tokens preserve localized contact.
During pretraining, the model masks portions of both modalities and reconstructs them. The encoder progressively merges 512, 256 and 64 point groups, then cross-attends the final geometric tokens with 19 tactile tokens. The decoder is discarded after pretraining; the encoder remains frozen while Proximal Policy Optimization learns the downstream controller. Architecture and protocol.
The human dataset uses a head-mounted Intel RealSense D435i and a WiseGlove with 19 tactile channels, synchronized at 30 Hz. Three contributors provide 565, 565 and 560 sequences. Pretraining runs on two RTX 4080 GPUs. The downstream simulated Shadow Hand receives a 4,096-point partial cloud, 19 binary contacts and 48 proprioceptive values at each control step.
Results and the percentage-point check
Faucet rotation uses three training faucet geometries and tests scales of 0.9 and 1.1 as unseen conditions. OccluDex reports 99.9% success on seen faucets and 86.6% on the unseen scales. Tabletop reorientation trains on ten YCB objects and holds out six; success is 75.1% on seen objects and 83.3% on held-out objects.
The paper's abstract says OccluDex is 8.3% higher on seen cases and 12.6% higher on unseen cases than the strongest baselines. RoboSkin recalculated those statements from Table III. Averaging the two tasks gives OccluDex 87.5% seen success versus 79.2% for the strongest per-task baseline, a difference of 8.3 percentage points. For unseen cases, 84.95% versus 72.4% produces 12.55 points, which rounds to 12.6. These are absolute percentage-point gaps, not relative percentage gains.
The physical platform uses a right Shadow Hand with 19 force-sensitive resistors, a RealSense D435i and closed-loop control at 10 Hz. The tactile electronics sample at 1 kHz before binarization; depth arrives at 30 Hz. Two unseen objects are used for each task. Faucet rotation succeeds in 17/20 trials, while 180-degree tabletop reorientation succeeds in 14/20.
What this means for robot hands
RoboSkin analysis: the result is useful because it tests a failure caused by the robot's own embodiment. More external cameras can reduce occlusion in a laboratory, but a dexterous robot hand operating from an onboard view still needs local evidence when fingers cover the object. OccluDex uses touch as that local evidence without asking the tactile stream to reconstruct the entire shape.
The ablation strengthens this interpretation. Removing touch drops faucet success by 45.7 points on seen instances and 48.0 points on unseen scales under the authors' ablation protocol. Removing the point cloud also hurts, but less consistently. The two modalities are complementary rather than interchangeable, which fits the broader visuo-tactile learning design problem.
Limitations and availability
OccluDex is an arXiv v1 manuscript submitted to IEEE, and RoboSkin.ai has not reproduced it. The physical study is 40 trials on one Shadow Hand setup. The real camera pose is matched to simulation, no point-cloud perturbation or dynamics randomization is used, and the policy sees binary contacts rather than rich tactile images or forces. The authors also note that training the full encoder jointly with PPO was infeasible on their hardware; the from-scratch comparison therefore uses a simplified architecture.
At verification time, the paper and arXiv record did not link a project page, public code, dataset, trained model or implementation license. The arXiv article is CC BY 4.0, but that does not establish reuse rights for the unreleased dataset or software.


