Attention Deficit · research story
Robotics · video · physical work

Can Robots Learn Physical Work by Watching People?

Dyna-2 used more than one million hours of first-person human video before post-training on robot demonstrations. The reported results support transfer from human video. Each evaluated task also used robot-specific data.

Sources: Dyna Robotics' technical report, Open X-Embodiment, EgoScale, Physical Intelligence's π0 policy, embodiment research, and Associated Press reporting on worker demonstrations.

Dyna-2 robot manipulation examples
Dyna-2 research footage. Image: Dyna Robotics. Technical report
Zoom out

Robot training data is expensive to collect.

The data ecosystem

Robot actions cost more to collect than web data.

TEXT + IMAGESThe public web

Large text and image collections describe objects and activities.

ROBOT ACTIONSpecialized collection

A robot, task setup, operator, sensors, and resets produce each episode.

HUMAN VIDEOPretraining data

First-person footage shows hand-object interactions and task sequences.

Human video can teach patterns of motion and object interaction before robot-specific training.
Context: Dyna-2; Open X-Embodiment; EgoScale; Physical Intelligence π0.

Robot demonstrations associate visual observations with body-specific controls.

Language models can learn broad patterns from copied digital artifacts. Robot data requires hardware time, teleoperation, object resets, safety supervision, and a specific body. Researchers pool data from many robots, use internet-scale visual pretraining, simulate experience, and learn physical structure from human video.

Scaling methods

Robot and human datasets provide different training signals.

OPEN X-EMBODIMENT1M+

robot episodes from 22 embodiments, 33 labs, and more than 500 skills

Unit: robot episodes
EGOSCALE20,854

hours of action-labeled egocentric human video for cross-embodiment transfer

Unit: human-video hours
DYNA-21M+

hours of egocentric human video used for model pretraining

Unit: human-video hours
1M hours ≈ 170 waking years
Dataset scale depends on embodiment, sensor, label, and measurement unit.
Sources: Google DeepMind Open X-Embodiment; NVIDIA EgoScale; Dyna-2.

Robot episodes and human-video hours measure different data.

Open X-Embodiment pools more than one million robot episodes across 22 embodiments. EgoScale organizes 20,854 hours of human egocentric video with action structure. Dyna reports more than one million hours of human video, roughly 170 years of waking time. The reported totals use robot episodes, labeled video hours, or raw pretraining hours, which prevents direct comparison.

Zoom in

Dyna-2 predicts future video frames before robot post-training.

The learning process

Human-video pretraining precedes robot adaptation.

01 · WATCHFirst-person human video

Hands and objects move through task sequences.

02 · PREDICTA plausible future frame

The model learns how actions change a scene.

03 · ADAPTLimited robot demonstrations

Up to ten hours per task connect visual knowledge to robot action.

04 · ACTPredict robot motion

The policy generates an action plan conditioned on the robot scene.

Human video supplies visual dynamics. Robot demonstrations connect the representation to robot controls.
Source: Dyna-2 technical report. Diagram simplifies the reported pretraining and post-training process.

Future-frame prediction supports later robot training.

Dyna trains a world-action model on video clips and action-conditioned future frames. During robot post-training, the action representation becomes robot motion. The company evaluates 14 tasks, each using no more than ten hours of robot data. The training process combines large-scale human-video pretraining with smaller task-specific robot datasets.

Evidence

Aggregate scores rise with human-video scale. Per-task scores vary.

The aggregate result

Average robot score increases with pretraining scale.

53%

aggregate score after 1M hours of human-video pretraining, across 14 robot tasks

The aggregate score increased from 20% at 1,000 hours to 53% at one million hours.
Source: Dyna-2 technical report. Aggregate normalized scores: 20%, 28%, 45%, 53%.

The model pretrained on one million hours led on nine of fourteen tasks.

Dyna reports aggregate normalized scores of 20%, 28%, 45%, and 53% as human-video pretraining increases from 1,000 to one million hours. The company held the robot post-training budget fixed within each task and ran ten evaluation trials per task, except for one language-conditioned evaluation with twelve.

Per-task results

Scores vary with pretraining scale.

LOCKBOX · USE KEYIncrease at 1M

Reported success: 0, 0, 0, 90 percent.

ROPE · TIENon-monotonic

Reported success: 0, 40, 90, 40 percent.

BOTTLE · CAPLow-data transfer

Thirteen minutes of teleoperation; reported success: 10, 10, 40, 50 percent.

Per-task results include increases, regressions, and flat regions across pretraining scales.
Source: Dyna-2 per-task results. Ten evaluation trials make each success count a ten-point step.

Per-task results show changes hidden by the aggregate score.

The lockbox task shows a threshold-like jump only at the largest scale. Rope tying peaks at 100,000 hours and declines at one million. Bottle capping improves with roughly ten minutes of robot teleoperation. Ten trials make each success worth ten percentage points, so task-level estimates carry wide uncertainty.

Embodiment

Human hands and robot grippers produce different action data.

Cross-embodiment transfer

Human video leaves a robot-specific control problem.

HUMAN VIEWHuman hand and tactile feedback
  • many degrees of freedom
  • tactile feedback
  • familiar tools
  • unrecorded effort
EMBODIMENT
GAP
ROBOT VIEWRobot sensors and controls
  • different geometry
  • limited sensing
  • motor latency
  • collision constraints
Robot demonstrations map visual task knowledge onto robot geometry and controls.
Sources: Dyna-2; Open X-Embodiment; recent research on the embodiment gap.

Video prediction provides incomplete evidence for motor control.

A head-mounted camera records object movement while forces, tactile cues, joint constraints, and corrective micro-actions remain partly unobserved. A robot also has different geometry and sensing. Recent embodiment research treats the remaining mismatch as a central transfer problem.

Beyond video

New datasets record physical signals that a camera misses.

VIDEOWhat moved?

First-person footage records objects, hands, and task order.

pixels · audio · scene context
BODY MOTIONHow did the person move?

Pose, gaze, and inertial sensors recover movement outside the image.

joint pose · gaze · acceleration
CONTACTHow hard did they push?

Force and tactile sensing expose pressure, slip, and contact changes.

force · touch · grip
ROBOT ROLLOUTCan its body execute?

Robot trials connect the recorded skill to motors and safety limits.

actions · state · recovery
Wearable sensing can reduce missing physical information while increasing the amount of worker data collected.
Sources: Meta Ego-Exo4D; Feel the Force; Associated Press reporting on instrumented worker demonstrations.

Robot-training collection is expanding beyond head-mounted cameras.

Meta's Ego-Exo4D pairs first-person video with external cameras, body and hand pose, eye gaze, and inertial measurements across skilled activities. Feel the Force studies tactile behavior for force-sensitive manipulation. Associated Press reporting describes commercial collection rigs that capture worker joint angles and applied force. Richer sensing addresses information missing from ordinary video and creates a more detailed record of a worker's technique.

People in the dataset

First-person video captures work and private information.

LABORWhose craft becomes robot capability?

Workers demonstrate handling, sequencing, recovery, and small adaptations rarely documented in a procedure manual.

consent · compensation · credit · bargaining power
PRIVACYWhat else appears in first-person video?

Faces, screens, conversations, locations, routines, and sensitive attributes can be visible or inferred even when the target label is only an action.

collection · filtering · retention · downstream access
First-person datasets contain worker demonstrations, information about the wearer, and footage of people nearby.
Sources: Associated Press reporting on worker demonstrations; EgoPrivacy research on first-person video.

First-person video creates labor and privacy obligations.

Associated Press reporting documents workers filmed while performing physical tasks to help train robots. EgoPrivacy shows that models can infer private attributes about the camera wearer. First-person footage can also record people nearby. Collection practices therefore affect workers and people visible in the recordings.

Evidence needed

Independent replication would strengthen the scaling claim.

01Independent tasksEvaluate outside the company's chosen suite and lab setup.
02Repeated runsReport uncertainty across seeds, checkpoints, and more trials.
03Data ablationsSeparate scale from diversity, quality, duplication, and relevance.
04Governed collectionDocument consent, filtering, provenance, access, and worker terms.
Evaluation across independent tasks and labs would test when human-video pretraining improves robot performance.
Research synthesis from Dyna-2 and adjacent robotics, embodiment, privacy, and labor sources.

Independent replication would test generalization.

The reported gain from one million hours of human video justifies further testing. Independent studies should include transparent data accounting, uncertainty estimates, diverse robot bodies and tasks, and governance records for human footage.