Workhorse 🐴

Learning Robust Whole‑Body Humanoid Loco‑Manipulation from Human Data

Songbo Hu*, Qiayuan Liao*, Yufeng Chi, Kevin Zakka,
Yakun Sophia Shao, Pieter Abbeel, Koushil Sreenath

University of California, Berkeley
*Equal contribution

Video Paper arXiv Code

Abstract

Humanoid robots still struggle to plan contact-rich whole-body manipulation from egocentric RGB and proprioception. Workhorse learns such manipulation from robot-free human demonstrations. A visual planner predicts five-link targets: the poses of the torso, both wrists, and both feet. A reinforcement-learning whole-body tracker follows them on the robot. Both policies train separately on the same recorded human poses, without retargeting. We augment the training data of each policy to imitate the errors that the other makes at deployment. On a real Unitree G1, Workhorse sorts boxes with its hands and a kick, catches a thrown box, and topples and climbs a suitcase. During box sorting, we show recoveries after a person pushes the robot or takes the box away. In a simulated copy of the demonstration room, the system completes box sorting in 77% of episodes, and in 64% under 40 N·s pushes. With both policies retrained from the same demonstrations, a simulated second humanoid completes box sorting in 83% of episodes without pushes.

Human data

A demonstrator carrying a box, wearing a chest camera on a harness, a tracker on the
                back of each hand and one on each shoe, with four base stations set around them.

The setup. Five wearable trackers (chest, both hands, both feet) and one chest camera. Nothing on the head, no markers on the objects.

What is recorded. The five wearable-tracker poses at 60 Hz and the camera's own video, put on one clock by three claps at the start of each session. One demonstrator recorded 2.2 h for box sorting, 1.2 h for the suitcase task and 13 min for box catching.

From human to humanoid

Alignment

Five wearable-tracker poses, one fixed offset per link. Nothing is retargeted.

Whole-body tracker

The whole-body tracker follows those same five-link targets, the markers, here in simulation. For the H2, only the five mounts and the robot model change, and both policies are retrained on the same demonstrations.

G1 Drag to turn
H2 Drag to turn

Visual planner

Input: four images from the robot's own camera and nine five-link samples from the last second. Output: a chunk of five-link targets 1.16 s long, drawn as the skeleton ahead of the robot. The flow-matching planner replans at 5 Hz, and at each step the whole-body tracker reads the next 0.2 s of the chunk.
Robot's camera

Results

Three tasks on the real G1. The robot acts from its own camera and proprioception alone, with no motion capture or teleoperation. Each task has its own planner and tracker; the interface, the training recipe and the deployment stack are shared.

Training

Whole-body tracker

On the robot, estimation, planning and tracking errors make the command drift between replans, and it jumps back at each replan. Training reproduces this sawtooth: a rigid drift of up to 10 cm horizontally and 10° in yaw, reset after a random 0.2 to 1 s. The reward follows the drifted command. At step t, the tracker reads the next 0.2 s of the command, the blue dots.
t time replan demo drifted command command window

Visual planner

Trained on a person, deployed on a robot. Each recorded image is edited once, offline: the demonstrator's body segmented out, the hole inpainted, the robot rendered in its place.
recorded edited Drag the divider
The tracker never follows a chunk exactly, so at deployment the planner sees a history that no demonstration contains. In training, the drifted history is the demonstrated history moved by a random walk, going back in time from the observation time t0. The chunk stays the clean demonstration, so the planner learns to recover from the drift.
t₀ time demo drifted history chunk: clean demo
The drift also moves the camera, which is fixed to the torso. Each image of the history is warped to the view the camera would have had after the drift. The images then agree with the drifted five-link history. The figure below runs the same warp live.
0° pitch   0° yaw Drag to turn the camera

Deployment

The training data was synchronized, but the streams on the robot are not. So every value carries the time it describes. The chunk starts at t0, when the newest image was captured, not when the chunk arrived. Actions already past (hollow) are skipped. The tracker reads the next 0.2 s of the chunk by timestamp, and each new chunk continues the previous one where they overlap.
camera delay + inference t t₀ image captured chunk arrives five-link hist. image hist. chunk action  k  at  t₀ + k Δt overlap with previous chunk tracker reads at t

Simulation evaluation

To test a policy before it goes on the robot, we scan the room into a Gaussian splat and evaluate there.

Success rate on box sorting without pushes, 100 episodes each. G1: 77%. H2: 83%. Under 40 N·s pushes, the G1 succeeds in 64%.

G1
H2
Success on the G1 when one part changes.
No pushesPushes
Full system77%64%
Joint-angle command from IK54%42%
Targets relative to each link5%4%
No gravity direction in the history1%1%
No planner augmentation20%0%
No command augmentation76%52%

Failure cases