OTRetarget: Joint Robot and Object Motion Retargeting via Optimal Transport
Abstract
Transferring human motion to humanoid robots requires adapting the demonstrated motion to the robot morphology while preserving interactions with the environment. This is particularly challenging for loco-manipulation tasks, where contacts with the ground and manipulated objects must remain consistent despite differences in body proportions. Yet, skeletal motion alone does not fully describe these interactions, and fixing object trajectories limits the adaptation to a new embodiment. In this paper, we introduce OTR ETARGET, a unified approach to jointly retarget robot and multi-object motion from human demonstrations. Our approach represents surface interactions through signed distances, closest surface points, and relative directions, and uses entropic optimal transport to transfer these quantities across human, robot, and object geometries. We incorporate the resulting interaction targets into a constrained inverse kinematics formulation that balances contact preservation with motion style and jointly optimizes robot and object poses at each frame. This formulation accommodates robot-object and object-object interactions without rescaling the scene or the demonstration. We validate the proposed approach on OMOMO, where it achieves a robot- object interaction Jaccard score of 87% and a depth error of 8.7 mm, compared with 28% and 29.3 mm for OmniRetarget. Finally, we demonstrate transfer to a physical G1 humanoid using whole-body policies trained with reinforcement learning on the retargeted references, across motions including two-handed box pick-and-place onto a table.
Community
OTRetarget retargets human demonstrations to a humanoid robot together with the objects it manipulates. Contacts are described by signed distances, closest surface points and relative directions, transferred across human, robot and object geometries with entropic optimal transport, then enforced in a constrained IK that optimizes robot and object poses jointly at each frame, without rescaling the scene or the demonstration. On OMOMO, it reaches a robot-object interaction Jaccard of 87% and a depth error of 8.7 mm, against 28% and 29.3 mm for OmniRetarget. Whole-body RL policies trained on the retargeted references transfer to a real Unitree G1, including two-handed box pick-and-place onto a table.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- HOI-Retarget: Contact-Centric Retargeting for Human-Object Interaction (2026)
- Unified Motion Retargeting for Humanoids with Learned Point Cloud Correspondence (2026)
- DexWeave: Learning Dexterous Humanoid Loco-Manipulation from Human Demonstrations (2026)
- Weave: Learning Whole-Body Dexterous Loco-Manipulation from Human-Object Interactions (2026)
- Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy (2026)
- Learning Scene-Aware Humanoid Locomotion through 3D Clutter from Immersive Human Demonstrations (2026)
- Gated Residual Body-Hand Coordination for Whole-Body Humanoid Teleoperation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.36602 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper