Can Robots Learn Physical Work by Watching People?
Dyna-2 used more than one million hours of first-person human video before post-training on robot demonstrations. The reported results support transfer from human video. Each evaluated task also used robot-specific data.
Sources: Dyna Robotics' technical report, Open X-Embodiment, EgoScale, Physical Intelligence's π0 policy, embodiment research, and Associated Press reporting on worker demonstrations.
Robot training data is expensive to collect.
Robot actions cost more to collect than web data.
Large text and image collections describe objects and activities.
A robot, task setup, operator, sensors, and resets produce each episode.
First-person footage shows hand-object interactions and task sequences.
Robot demonstrations associate visual observations with body-specific controls.
Language models can learn broad patterns from copied digital artifacts. Robot data requires hardware time, teleoperation, object resets, safety supervision, and a specific body. Researchers pool data from many robots, use internet-scale visual pretraining, simulate experience, and learn physical structure from human video.
Robot and human datasets provide different training signals.
robot episodes from 22 embodiments, 33 labs, and more than 500 skills
Unit: robot episodeshours of action-labeled egocentric human video for cross-embodiment transfer
Unit: human-video hourshours of egocentric human video used for model pretraining
Unit: human-video hoursRobot episodes and human-video hours measure different data.
Open X-Embodiment pools more than one million robot episodes across 22 embodiments. EgoScale organizes 20,854 hours of human egocentric video with action structure. Dyna reports more than one million hours of human video, roughly 170 years of waking time. The reported totals use robot episodes, labeled video hours, or raw pretraining hours, which prevents direct comparison.
Dyna-2 predicts future video frames before robot post-training.
Human-video pretraining precedes robot adaptation.
Hands and objects move through task sequences.
The model learns how actions change a scene.
Up to ten hours per task connect visual knowledge to robot action.
The policy generates an action plan conditioned on the robot scene.
Future-frame prediction supports later robot training.
Dyna trains a world-action model on video clips and action-conditioned future frames. During robot post-training, the action representation becomes robot motion. The company evaluates 14 tasks, each using no more than ten hours of robot data. The training process combines large-scale human-video pretraining with smaller task-specific robot datasets.
Aggregate scores rise with human-video scale. Per-task scores vary.
Average robot score increases with pretraining scale.
aggregate score after 1M hours of human-video pretraining, across 14 robot tasks
The model pretrained on one million hours led on nine of fourteen tasks.
Dyna reports aggregate normalized scores of 20%, 28%, 45%, and 53% as human-video pretraining increases from 1,000 to one million hours. The company held the robot post-training budget fixed within each task and ran ten evaluation trials per task, except for one language-conditioned evaluation with twelve.
Scores vary with pretraining scale.
Reported success: 0, 0, 0, 90 percent.
Reported success: 0, 40, 90, 40 percent.
Thirteen minutes of teleoperation; reported success: 10, 10, 40, 50 percent.
Per-task results show changes hidden by the aggregate score.
The lockbox task shows a threshold-like jump only at the largest scale. Rope tying peaks at 100,000 hours and declines at one million. Bottle capping improves with roughly ten minutes of robot teleoperation. Ten trials make each success worth ten percentage points, so task-level estimates carry wide uncertainty.
Human hands and robot grippers produce different action data.
Human video leaves a robot-specific control problem.
- many degrees of freedom
- tactile feedback
- familiar tools
- unrecorded effort
GAP
- different geometry
- limited sensing
- motor latency
- collision constraints
Video prediction provides incomplete evidence for motor control.
A head-mounted camera records object movement while forces, tactile cues, joint constraints, and corrective micro-actions remain partly unobserved. A robot also has different geometry and sensing. Recent embodiment research treats the remaining mismatch as a central transfer problem.
New datasets record physical signals that a camera misses.
First-person footage records objects, hands, and task order.
pixels · audio · scene contextPose, gaze, and inertial sensors recover movement outside the image.
joint pose · gaze · accelerationForce and tactile sensing expose pressure, slip, and contact changes.
force · touch · gripRobot trials connect the recorded skill to motors and safety limits.
actions · state · recoveryRobot-training collection is expanding beyond head-mounted cameras.
Meta's Ego-Exo4D pairs first-person video with external cameras, body and hand pose, eye gaze, and inertial measurements across skilled activities. Feel the Force studies tactile behavior for force-sensitive manipulation. Associated Press reporting describes commercial collection rigs that capture worker joint angles and applied force. Richer sensing addresses information missing from ordinary video and creates a more detailed record of a worker's technique.
First-person video captures work and private information.
Workers demonstrate handling, sequencing, recovery, and small adaptations rarely documented in a procedure manual.
Faces, screens, conversations, locations, routines, and sensitive attributes can be visible or inferred even when the target label is only an action.
First-person video creates labor and privacy obligations.
Associated Press reporting documents workers filmed while performing physical tasks to help train robots. EgoPrivacy shows that models can infer private attributes about the camera wearer. First-person footage can also record people nearby. Collection practices therefore affect workers and people visible in the recordings.
Independent replication would strengthen the scaling claim.
Independent replication would test generalization.
The reported gain from one million hours of human video justifies further testing. Independent studies should include transparent data accounting, uncertainty estimates, diverse robot bodies and tasks, and governance records for human footage.