Abstract
Humanoid robots still struggle to plan contact-rich whole-body manipulation from egocentric RGB and proprioception. Workhorse learns such manipulation from robot-free human demonstrations. A visual planner predicts five-link targets: the poses of the torso, both wrists, and both feet. A reinforcement-learning whole-body tracker follows them on the robot. Both policies train separately on the same recorded human poses, without retargeting. We augment the training data of each policy to imitate the errors that the other makes at deployment. On a real Unitree G1, Workhorse sorts boxes with its hands and a kick, catches a thrown box, and topples and climbs a suitcase. During box sorting, we show recoveries after a person pushes the robot or takes the box away. In a simulated copy of the demonstration room, the system completes box sorting in 77% of episodes, and in 64% under 40 N·s pushes. With both policies retrained from the same demonstrations, a simulated second humanoid completes box sorting in 83% of episodes without pushes.
Human data
The setup. Five wearable trackers (chest, both hands, both feet) and one chest camera. Nothing on the head, no markers on the objects.
What is recorded. The five wearable-tracker poses at 60 Hz and the camera's own video, put on one clock by three claps at the start of each session. One demonstrator recorded 2.2 h for box sorting, 1.2 h for the suitcase task and 13 min for box catching.
From human to humanoid
Alignment
Whole-body tracker
The whole-body tracker follows those same five-link targets, the markers, here in simulation. For the H2, only the five mounts and the robot model change, and both policies are retrained on the same demonstrations.
Visual planner
Results
Three tasks on the real G1. The robot acts from its own camera and proprioception alone, with no motion capture or teleoperation. Each task has its own planner and tracker; the interface, the training recipe and the deployment stack are shared.
Training
Whole-body tracker
Visual planner
Deployment
Simulation evaluation
To test a policy before it goes on the robot, we scan the room into a Gaussian splat and evaluate there.
Success rate on box sorting without pushes, 100 episodes each. G1: 77%. H2: 83%. Under 40 N·s pushes, the G1 succeeds in 64%.
| No pushes | Pushes | |
|---|---|---|
| Full system | 77% | 64% |
| Joint-angle command from IK | 54% | 42% |
| Targets relative to each link | 5% | 4% |
| No gravity direction in the history | 1% | 1% |
| No planner augmentation | 20% | 0% |
| No command augmentation | 76% | 52% |