Everything K-culture — comebacks to K-beauty, straight to your inboxGet it in your inbox

METAL MEDIA

Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data

arXiv:2608.025802026-08-02

Researchers turned first-person human manipulation videos into 18,561 hours of training data for 15 robot types and tested whether it actually helps robot policies generalize

Training robots to manipulate objects well needs lots of varied demonstrations, but collecting them directly on real robots is slow and expensive. Ego2Robot is a pipeline that converts egocentric human videos into robot training data through action alignment, visual alignment, and quality filtering, producing 18,561 hours of data across 15 robot arm types. Mixing this synthetic data with real robot data during pretraining consistently improved success rates under unseen visual, layout, embodiment, and language conditions, with gains confirmed on a real robot as well.

METAL MEDIA explanatory visual

Ego2Robot pipeline: from human video to robot training data

Evidence statusMeasured results reported

  1. Input: egocentric manipulation videoAbout 1,940 hours of first-person human hand manipulation footage from four sources: ANT, EgoDex, ViTRA, EgoVerse
  2. Action alignment21 hand keypoints per frame are converted into smooth robot end-effector trajectories, with speed matched to robot teleoperation pace
  3. Visual alignmentSAM 3 segments and removes the human arm via inpainting; a robot base pose is solved via inverse kinematics and a rendered robot arm (one of 15 types) is composited into the scene
  4. Quality curationThree levels of filtering remove IK failures, self-collisions, statistical outliers, and semantically inconsistent clips flagged by a vision-language model
  5. Output: 18,561 hours of training dataParallel data generated for 15 robot morphologies, mixed with real robot data to pretrain a vision-language-action model
An explanatory diagram made by METAL MEDIA, not a figure supplied by the paper's authors.

What they did

  1. The team collected about 1,940 hours of first-person human hand manipulation videos from four sources (ANT, EgoDex, ViTRA, EgoVerse), then built a pipeline with three stages: action alignment (converting 21 hand keypoints into robot end-effector trajectories), visual alignment (removing human arms via SAM 3 segmentation and inpainting, then compositing a rendered robot arm using inverse kinematics), and multi-level quality curation to filter bad frames and trajectories.
  2. They ran this pipeline separately for 15 different robot arm morphologies (Panda, UR5e, xArm7, and 12 others), turning each source video into 15 parallel training streams and yielding a total of 18,561 hours of synthesized robot data, described as the largest ego-to-robot dataset so far.
  3. To measure generalization precisely, the authors extended the RoboTwin2.0 benchmark so that four factors -- visual appearance (background, lighting, robot color), scene layout (table height, clutter, camera offset), embodiment (swapping in different robot arms), and task semantics (unseen objects, paraphrased instructions) -- could each be tested independently instead of all mixed together.
  4. Compared with pretraining on real robot data alone (about 6,565 hours from DROID, AgibotWorld, and InternData), mixing the synthesized data with real data at a 1:1 ratio raised success rate on RoboTwin's randomized setting to 53.5% (+2.6 points), and produced gains across background (+4), lighting (+8), robot color (+6), camera offset (+6), unseen objects (+11 at a 3:1 ratio), and paraphrased instructions (reaching 69% at 1:1).
  5. On a real ARX ACone robot across five tasks (putting fruit in a basket, placing blocks in a drawer, folding a towel, sweeping trash, inserting a screw), the model trained with both teleoperated demonstrations and the synthesized ego-derived data outperformed the robot-only model on every task, with the biggest gains on placing blocks (+14 points) and inserting a screw (+13 points).
Figure 1: Ego2Robot pipeline. Converting egocentric video into 18,561h robot training data across 15 morphologies through action alignment, visual alignment, and quality curation.
Figure 1: Ego2Robot pipeline. Converting egocentric video into 18,561h robot training data across 15 morphologies through action alignment, visual alignment, and quality curation.
Table 1: Main results. Success rates (%) on RoboTwin2.0 and EBench. Green = gain >5% vs. Robot-only.
PretrainingRoboTwinPer-Dimension (RoboTwin)EBench
CleanRandVisualSceneEmbodyTaskAvg
Robot-only62.250.961.452.923.846.239.6
Ego2R+Robot (1:3)61.4 –0.851.0 +0.161.2 –0.252.5 –0.421.9 –1.949.5 +3.347.4 +7.8
Ego2R+Robot (3:1)64.1 +1.949.2 –1.762.7 +1.354.3 +1.428.2 +4.451.6 +5.451.7 +12.1
Ego2R+Robot (1:1)68.1 +5.953.5 +2.667.3 +5.956.9 +4.027.2 +3.454.1 +7.949.8 +10.2
Figure 2: Evaluation framework. Four generalization dimensions with 12 evaluation settings. Dashed borders: settings decoupled from bundled randomization for independent testing. Solid borders: newly introduced evaluation axes. Gray border: external benchmark (EBench).
Figure 2: Evaluation framework. Four generalization dimensions with 12 evaluation settings. Dashed borders: settings decoupled from bundled randomization for independent testing. Solid borders: newly introduced evaluation axes. Gray border: external benchmark (EBench).
Table 2: Per-perturbation breakdown. Change vs. Robot-only. Green = gain >5%.
PretrainingVisualSceneEmbodimentTask
BGLightColorHeightClutterCameraARXUR5FrankaObjLang
Robot-only66.658.259.460.148.350.444.120.27.029.363.1
Ego2R+Robot (1:3)65.0 –1.658.3 +0.160.3 +0.958.6 –1.549.3 +1.049.6 –0.843.7 –0.417.6 –2.64.5 –2.536.8 +7.562.2 –0.9
Ego2R+Robot (3:1)65.5 –1.160.9 +2.761.8 +2.462.0 +1.949.2 +0.951.6 +1.247.6 +3.531.4 +11.25.6 –1.440.0 +10.763.1 +0.0
Ego2R+Robot (1:1)70.3 +3.765.8 +7.665.8 +6.462.4 +2.352.0 +3.756.3 +5.951.2 +7.125.0 +4.85.3 –1.739.6 +10.368.5 +5.4
Figure 3: Pipeline value and embodiment scaling. Success rate on RoboTwin Randomized.
Figure 3: Pipeline value and embodiment scaling. Success rate on RoboTwin Randomized.
Table 3: Supported robot morphologies.
RobotDOFGripper (mm)Reach (m)
Panda70–801.272
Kinova Gen370–851.337
IIWA70–851.411
Sawyer714–791.420
FR370–801.272
xArm770–851.290
UR5e60–851.236
UR10e60–851.627
Jaco60–1251.200
ViperX615–870.911
WidowX611–550.787
ARX-L560–880.855
Piper60–700.883
YAM64–750.866
Aloha-Agilex67–1020.853
Figure 4: Real robot results. Success rates (%) on five tasks on the ARX ACone platform.
Figure 4: Real robot results. Success rates (%) on five tasks on the ARX ACone platform.
Table 4: Real robot scoring.
TaskStep 1Step 2Step 3Step 4
Put Fruits33.333.333.3
Put Blocks25252525
Fold Towel5050
Sweep Trash25252525
Insert Screw25252525
Figure 5: Real robot rollouts. Key frames from five evaluation tasks on the ARX ACone platform.
Figure 5: Real robot rollouts. Key frames from five evaluation tasks on the ARX ACone platform.
Table 5: Per-task results on RoboTwin including Pi0.5 (success rate %).
TaskPi0.5Robot-onlyEgo2R (1:3)Ego2R (3:1)Ego2R (1:1)
CleanRandCleanRandCleanRandCleanRandCleanRand
adjust_bottle62.018.094.077.092.079.0100.072.089.091.0
beat_block_hammer48.016.063.027.064.056.057.018.065.071.0
blocks_ranking_rgb74.016.056.060.064.063.049.043.078.060.0
blocks_ranking_size32.00.030.022.046.021.043.07.052.016.0
click_alarmclock44.056.096.093.0100.087.0100.085.0100.0100.0
click_bell32.040.0100.095.0100.094.0100.093.0100.0100.0
dump_bin_bigbin72.070.054.073.072.069.081.076.086.072.0
grab_roller94.046.098.071.068.048.091.064.089.063.0
handover_block16.00.05.02.018.04.039.09.036.09.0
handover_mic26.04.069.013.086.034.083.017.097.023.0
hanging_mug10.04.014.016.014.09.012.09.010.010.0
lift_pot8.02.093.028.084.014.095.044.093.044.0
move_can_pot28.00.047.050.030.040.042.085.047.065.0
move_pillbottle_pad60.044.063.060.068.069.041.067.074.082.0
move_playingcard_away90.052.074.058.088.092.088.042.092.063.0
move_stapler_pad22.06.023.019.034.021.030.013.043.031.0
open_laptop68.020.077.061.074.056.064.046.075.058.0
open_microwave26.08.077.059.042.029.037.032.050.025.0
pick_diverse_bottles56.018.060.037.058.053.071.050.070.050.0
pick_dual_bottles82.014.091.059.080.077.083.055.097.054.0
place_a2b_left60.010.040.055.058.042.054.048.073.044.0
place_a2b_right58.014.042.049.052.041.053.051.064.053.0
place_bread_basket68.044.075.055.072.064.076.059.079.062.0
place_bread_skillet86.046.076.041.076.050.080.061.087.053.0
place_burger_fries90.078.096.083.098.081.097.088.096.088.0
place_can_basket40.00.049.025.038.09.034.015.034.016.0
place_cans_plasticbox94.074.090.068.054.068.097.059.070.059.0
place_container_plate96.058.086.076.092.079.093.073.093.081.0
place_dual_shoes54.012.050.030.034.027.035.017.035.019.0
Figure 6: 15 supported robot morphologies. 3D models of all 15 robots.
Figure 6: 15 supported robot morphologies. 3D models of all 15 robots.

Findings

  • Mixing Ego2R-synthesized data with real robot data at a 1:1 ratio reached 53.5% success on RoboTwin's randomized setting, 2.6 points above robot-only pretraining, while still keeping 68.1% on the clean setting.
  • At the 1:1 mixing ratio, per-factor visual improvements were background +4, lighting +8, robot color +6, and camera offset +6 compared to robot-only pretraining.
  • Unseen-object generalization improved from 29% to 40% (+11 at a 3:1 ratio), and success on paraphrased instructions reached 69% at the 1:1 ratio.
  • In cross-embodiment tests swapping in different robot arms, ARX improved from 44% to 51% and UR5 peaked at 31% (3:1 ratio), while Franka Panda stayed below 7%.
  • On the real ARX ACone robot across five tasks, the model trained with both teleoperated demonstrations and the pipeline-converted egocentric data outperformed the robot-only model on every task, with the largest gains on placing blocks and inserting a screw.
Figure 7: Pretraining data composition. Slices show per-source sampling weights (training mix); panel titles give the total data volume. (a) Robot (∼6,565 h). (b) Ego2R (∼18,561 h).
Figure 7: Pretraining data composition. Slices show per-source sampling weights (training mix); panel titles give the total data volume. (a) Robot (∼6,565 h). (b) Ego2R (∼18,561 h).

Where it can be used

  • Teams that lack large amounts of real demonstration data for a new robot arm or new task could consider converting existing human manipulation footage into supplementary pretraining data.
  • Service robot settings where backgrounds, lighting, objects, and phrasing change frequently could consider using diverse human video as an additional training source to improve robustness.
  • Projects supporting multiple robot arm models at once could reference this approach of rendering a single human video into several robot embodiments in parallel.
Figure 8: Ego play → Ego2R synthesis on real-world scenes. Top: original egocentric human manipulation. Bottom: Ego2R pipeline output with ACone robot overlay.
Figure 8: Ego play → Ego2R synthesis on real-world scenes. Top: original egocentric human manipulation. Bottom: Ego2R pipeline output with ACone robot overlay.

Limits and open work

  • Hand motions are converted only into simple parallel-jaw gripper movements, so the approach does not directly extend to dexterous multi-fingered manipulation.
  • Visual synthesis relies on inpainting and depth-aware compositing, which can introduce artifacts under heavy occlusion or difficult lighting conditions.
  • Evaluation is confined to the task scope of RoboTwin2.0, so it remains untested whether the benefits hold for a broader range of tasks and robot configurations.
  • For robots with a large kinematic gap from the training embodiments, such as Franka Panda, cross-embodiment success stayed below 7%, showing the gains are not uniform across all robot types.
Figure 9: Comparison with Pi0.5 across RoboTwin settings.
Figure 9: Comparison with Pi0.5 across RoboTwin settings.

Why it matters

Collecting real robot demonstrations is expensive and slow, so being able to automatically convert the enormous existing supply of first-person human manipulation videos into usable robot training data could substantially cut data-collection costs. The fact that this specifically improves robustness to unseen backgrounds, robot types, objects, and phrasing matters directly for deploying robots in real, unpredictable environments.

Terms in this paper

  • Egocentric video · Video recorded from a person's own point of view, typically with a head- or body-mounted camera, showing their hands manipulating objects
  • VLA model (vision-language-action model) · An AI model that takes camera images and language instructions as input and directly outputs the robot actions to perform
  • Inverse kinematics (IK) · A method for computing the joint angles a robot arm needs so its end-effector reaches a desired position and orientation
  • Out-of-distribution (OOD) generalization · The ability to keep performing well under backgrounds, objects, robot types, or phrasing that were not seen during training
  • End-effector (EEF) · The part at the tip of a robot arm, such as a gripper, that actually makes contact with objects

Original abstract (English)

Learning generalizable robot manipulation policies requires large-scale and diverse demonstration data. Egocentric human manipulation videos offer rich scene and task diversity, and prior work has shown that retargeting and rendering such videos into robot-format data can yield effective per-task policies at small scale. However, whether this approach can provide pretraining benefits for vision-language-action models at scale remains unexplored. We present Ego2Robot, a scalable pipeline that con

Authors · Ye Wang

Read on arXiv

Latest papers

All papers →

Latest from METAL MEDIA

Figures: Ye Wang et al., arXiv:2608.02580, arxiv-nonexclusive