
1Nanjing University
2China Mobile Research Institute
*Equal contribution · †Corresponding author
From Human Demonstrations to Robot Dexterity
Human Video
Reconstructed HOI
Robot Execution
High-quality demonstrations for dexterous robot manipulation are costly and difficult to collect, whereas monocular human videos provide a scalable source of diverse manipulation behaviors. However, transferring such demonstrations to dexterous robots remains challenging: monocular hand–object interaction (HOI) reconstruction often produces temporally unstable contacts and physically implausible interactions, while conventional retargeting methods struggle to preserve task-relevant contacts and local interaction geometry across different hand embodiments. We present C2Dex, a video-to-dexterous-manipulation framework that integrates contact-consistent HOI reconstruction with interaction-preserving retargeting. C2Dex aggregates noisy frame-wise contact observations in the canonical object space to recover stable object-side contacts, which guide trajectory-level optimization toward temporally coherent and physically plausible human HOI trajectories. It then transfers these stable contacts to the target dexterous hand, employs Laplacian interaction optimization to preserve the surrounding local hand–object geometry across embodiments, and refines the trajectory via residual reinforcement learning in simulation. Experiments on DexYCB and TACO show that C2Dex achieves end-to-end trajectory success rates of 57.78% and 26.67%, respectively, substantially outperforming the strongest baselines (17.78% and 10.00%) under identical evaluation criteria. Real-robot experiments further verify direct execution across diverse contact-rich manipulation tasks.
C2Dex converts a monocular human video into an executable dexterous manipulation trajectory through two tightly coupled modules. Contact-consistent HOI reconstruction recovers stable object-side contacts from noisy frame-wise observations and uses them to refine the human HOI trajectory. Contact-interaction-preserving retargeting transfers the reconstructed interaction to the target dexterous hand while preserving task-relevant contacts and local hand–object geometry.
Given a monocular video, we first estimate an initial human HOI trajectory and extract frame-wise contact observations between the hand and object, which are then stabilized and used to refine the hand motion.
Starting from a keypoint-based initialization, stable contact optimization transfers task-relevant contacts while Laplacian optimization preserves local hand–object geometry. Residual reinforcement learning then yields the executable trajectory.
We collect 24 human demonstration videos covering 8 contact-rich daily manipulation tasks. For each demonstration, C2Dex reconstructs the human HOI and generates the corresponding dexterous-hand–object trajectory, which is then directly replayed on a physical robot equipped with the Inspire dexterous hand — no per-task policy training and no teleoperation.
For each benchmark sequence we show the human demonstration video alongside the rollout of the residual policy in simulation.