H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models
A new benchmark checks whether AI-generated 'robot versions' of human demo videos actually preserve the real action, not just look nice
Robot training videos are scarce and expensive, while human videos of manipulating objects are abundant, so researchers are testing whether video generation AI can convert human demo videos into robot demo videos. Past evaluations only asked whether the generated video looked plausible, not whether the actual task, contact, and robot body type were correctly transferred. H2R-Bench scored 11 state-of-the-art video generation models on 120 human demonstration clips converted into two robot body types, and found that visual polish barely predicts whether the manipulation was actually transferred correctly.
METAL MEDIA explanatory visual
H2R-Bench evaluation pipeline
Evidence statusMeasured results reported
- Input: human demo videos120 egocentric human manipulation clips from EgoDex, evenly spread across six task families
- Condition: target robot embodimenteach source video is targeted toward either a parallel-jaw gripper or a dexterous hand, giving 240 test cases
- Generation: 11 video AI modelsvideo-conditioned models like Seedance 2.0, Wan2.7, Kling-V3 and frame-conditioned models like HunyuanVideo generate the robot video via their own native interfaces
- Scoring: five metrics M1-M5goal completion, action completion, functional contact, embodiment correctness, and video quality, each scored 0-4 by three MLLM judges
- Output: H2RCore scorethe five scores are combined into a 0-100 score, with contact and embodiment each weighted 30%
What they did
- The team selected 120 egocentric human manipulation clips from the EgoDex dataset and asked 11 video generation AI models to turn each into a robot video, once for a parallel-jaw gripper and once for a dexterous five-fingered robot hand, giving 240 test cases.
- Generated videos were scored on five dimensions: goal-state completion (M1), whether required actions occurred (M2), functional contact between robot and object (M3), whether the requested robot embodiment was correctly shown (M4), and general video quality (M5). Three large multimodal AI judges (Gemini, Qwen, GPT) scored the videos, and their scores were checked against human raters.
- Models that could take the full source video as input, Seedance 2.0, Wan2.7, and Kling-V3, ranked at the top; Seedance 2.0 scored highest overall on the combined H2RCore metric (0-100 scale) with 77.3 for the gripper and 84.6 for the dexterous hand.
- Video quality scores (M5) were nearly identical across models, ranging only 0.73 to 0.81, while overall transfer scores (H2RCore) ranged widely from 30.0 to 84.6, with weak correlation between the two (Spearman rho=0.14) meaning a pretty-looking video did not mean the robot task was actually completed.
- HunyuanVideo 1.5-I2V had the best video quality score but a low contact score, often leaving the human hand doing the work instead of the robot; Veo 3.1 produced plausible interactions but frequently generated the wrong end effector for the requested embodiment.

| Benchmark | Settings | Evaluation | ||||||
|---|---|---|---|---|---|---|---|---|
| I2V | RV | H2R | VQ | Goal | Action | Cont. | Emb. | |
| VBench | × | × | × | ✓ | × | × | × | × |
| WorldModelBench | ✓ | △ | × | ✓ | △ | △ | × | × |
| RBench | ✓ | ✓ | × | ✓ | ✓ | ✓ | × | △ |
| RoboWM-Bench | ✓ | ✓ | × | △ | ✓ | ✓ | × | × |
| RoboTrustBench | ✓ | ✓ | × | ✓ | ✓ | ✓ | × | △ |
| H2R-Bench (ours) | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |

| Model | Parallel-Jaw Gripper | Dexterous Hand | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Component Metrics | Aggregate | Component Metrics | Aggregate | |||||||||
| Goal Comp. | Action Comp. | Contact Transfer | Embod. Correct. | Video Quality | H2R Core | Goal Comp. | Action Comp. | Contact Transfer | Embod. Correct. | Video Quality | H2R Core | |
| Video-conditioned generation | ||||||||||||
| Seedance 2.0 | 0.725 | 0.813 | 0.776 | 0.768 | 0.793 | 77.3 | 0.744 | 0.832 | 0.855 | 0.911 | 0.799 | 84.6 |
| Wan2.7 | 0.706 | 0.791 | 0.766 | 0.772 | 0.796 | 76.5 | 0.718 | 0.804 | 0.835 | 0.910 | 0.795 | 83.1 |
| Kling-V3 | 0.710 | 0.807 | 0.751 | 0.707 | 0.798 | 74.5 | 0.707 | 0.800 | 0.819 | 0.885 | 0.802 | 81.7 |
| Frame-conditioned generation | ||||||||||||
| Mitty-EPIC14B | 0.581 | 0.668 | 0.598 | 0.585 | 0.732 | 61.5 | 0.587 | 0.684 | 0.598 | 0.392 | 0.732 | 56.1 |
| Veo 3.1 | 0.725 | 0.797 | 0.533 | 0.100 | 0.783 | 49.6 | 0.715 | 0.816 | 0.642 | 0.227 | 0.793 | 57.0 |
| Grok Imagine Video | 0.661 | 0.729 | 0.443 | 0.268 | 0.792 | 50.1 | 0.678 | 0.729 | 0.469 | 0.198 | 0.804 | 49.2 |
| LTX-2.3 | 0.473 | 0.545 | 0.292 | 0.012 | 0.773 | 32.1 | 0.520 | 0.592 | 0.377 | 0.132 | 0.780 | 39.8 |
| SkyReels-V3-R2V | 0.448 | 0.610 | 0.256 | 0.004 | 0.787 | 31.5 | 0.441 | 0.607 | 0.341 | 0.026 | 0.789 | 34.6 |
| Wan2.2 | 0.492 | 0.639 | 0.258 | 0.000 | 0.769 | 32.4 | 0.512 | 0.653 | 0.286 | 0.000 | 0.766 | 33.7 |
| LongCat | 0.472 | 0.583 | 0.243 | 0.000 | 0.790 | 31.0 | 0.428 | 0.545 | 0.298 | 0.020 | 0.793 | 32.0 |
| HunyuanVideo 1.5-I2V | 0.535 | 0.549 | 0.184 | 0.005 | 0.806 | 30.0 | 0.499 | 0.555 | 0.185 | 0.041 | 0.808 | 30.7 |

| Metric | Mean change (Hand − Gripper) | Hand higher (models) |
|---|---|---|
| Goal-State Completion (M1) | +0.002 | 6/11 |
| Action-Event Completion (M2) | +0.008 | 8/11 |
| Functional Contact Transfer (M3) | +0.055 | 11/11 |
| Embodiment Correctness (M4) | +0.047 | 8/11 |
| Video Quality (M5) | +0.004 | 8/11 |
| H2RCore (0–100) | +3.3 | 9/11 |
| Model | Ref. | Goal | Action | Contact | Embod. | Quality | Core |
|---|---|---|---|---|---|---|---|
| Parallel-Jaw Gripper | |||||||
| Kling-V3 | No | 0.710 | 0.807 | 0.751 | 0.707 | 0.798 | 74.5 |
| Yes | 0.673 | 0.760 | 0.577 | 0.479 | 0.779 | 61.0 | |
| Seedance 2.0 | No | 0.725 | 0.813 | 0.776 | 0.768 | 0.793 | 77.3 |
| Yes | 0.711 | 0.794 | 0.674 | 0.595 | 0.784 | 68.5 | |
| Wan2.7 | No | 0.706 | 0.791 | 0.766 | 0.772 | 0.796 | 76.5 |
| Yes | 0.714 | 0.811 | 0.871 | 0.875 | 0.781 | 83.1 | |
| Dexterous Hand | |||||||
| Kling-V3 | No | 0.707 | 0.800 | 0.819 | 0.885 | 0.802 | 81.7 |
| Yes | 0.629 | 0.734 | 0.772 | 0.728 | 0.799 | 73.4 | |
| Seedance 2.0 | No | 0.744 | 0.832 | 0.855 | 0.911 | 0.799 | 84.6 |
| Yes | 0.736 | 0.824 | 0.876 | 0.753 | 0.794 | 80.2 | |
| Wan2.7 | No | 0.718 | 0.804 | 0.835 | 0.910 | 0.795 | 83.1 |
| Yes | 0.712 | 0.819 | 0.905 | 0.873 | 0.789 | 84.2 |

| Element | Embodied-manipulation concern | H2R-Bench operationalization | Design origin |
|---|---|---|---|
| Task families: task-defining physical state changes | |||
| F1: Rigid rearrangement | Object pose, support, or containment | Place or transport a rigid object into the demonstrated spatial relation. | Rigid transport and placement tasks. |
| F2: Mechanism actuation | State of an articulated mechanism | Open, close, toggle, press, or rotate a task-relevant mechanism. | Interaction with articulated objects. |
| F3: Insertion and assembly | Connection, fit, or attachment relation | Establish or remove a constrained connection between entities. | Precision alignment and constrained contact. |
| F4: Deformable configuration | Non-rigid shape or configuration | Produce the demonstrated fold, bend, compression, or shape change. | Deformable-object manipulation. |
| F5: Bulk-material transfer | Distribution or containment of material | Pour, scoop, transfer, or mix material between regions or containers. | Many-particle and material-flow manipulation. |
| F6: Surface/material transformation | Local surface condition or material integrity | Produce a visible local change through wiping, spreading, cutting, peeling, or related interaction. | Tool-mediated local transformation. |
| Evaluation dimensions: evidence required for valid H2R generation | |||
| M1: Goal-State Completion | Was the demonstrated task state reached? | Score weighted predicates over the visible final state. | Task-success and outcome evaluation in robotic video benchmarks. |
| M2: Action-Event Completion | Were the required manipulation events shown? | Score completion of source-derived action events. | Action-completeness evaluation in robotic video benchmarks. |
| M3: Functional Contact Transfer | Did robot contact support the source-consistent object response? | Evaluate contact region, establishment, mode, temporal object response, and embodiment-compatible strategy. | Contact-mediated manipulation and H2R interaction grounding. |
| M4: Embodiment Correctness | Did the requested robot perform the manipulation? | Evaluate robot presence, human absence, embodiment category, end-effector subtype, and temporal structure. | Robot-structure and embodiment compliance. |
| M5: Video Quality | Is the generated video visually and temporally well formed? | Measure imaging quality, aesthetic quality, temporal stability, and motion smoothness. | Task-agnostic video-generation quality. |
| Stage | Output | Quality control |
|---|---|---|
| Source curation | Curated clips across six families | Identifiable manipulated entities, visible task-relevant state change, and sufficient interaction evidence. |
| Clip grounding | 32 ordered source frames per clip | First and last frames are retained; annotations identify initial state, contact, core transition, and final state when visible. |
| Structured annotation | Goals, events, milestones, interaction, and embodiment strategies | Strict JSON schema with canonical event identifiers and a fixed six-family taxonomy. |
| Completion and validation | Schema-complete case records | Only missing fields are completed; existing non-empty fields are preserved; schema and type mismatches are rejected. |
| Manual verification | All benchmark annotations | Every record is checked against its source video and corrected before evaluation. |
| Uncertainty accounting | Validity and ambiguity fields | Every record has a validity decision; notes document residual uncertainty without changing the official evaluation set. |
| Target | Input | M1 | M2 | M3 | M4 | M5 | Core |
|---|---|---|---|---|---|---|---|
| Parallel-Jaw Gripper | Video | 0.737 | 0.850 | 0.709 | 0.598 | 0.793 | 70.9 |
| 9 images | 0.744 | 0.848 | 0.284 | 0.104 | 0.783 | 43.4 | |
| Dexterous Hand | Video | 0.768 | 0.857 | 0.861 | 0.898 | 0.799 | 85.1 |
| 9 images | 0.753 | 0.856 | 0.313 | 0.079 | 0.782 | 43.7 |
| Model | Parallel-Jaw Gripper Core [95% CI] | Dexterous Hand Core [95% CI] |
|---|---|---|
| Kling-V3 | 74.5 [73.6, 75.3] | 81.7 [80.8, 82.6] |
| Seedance 2.0 | 77.3 [76.5, 78.1] | 84.6 [84.0, 85.3] |
| Wan2.7 | 76.5 [75.7, 77.4] | 83.1 [82.3, 83.9] |
| Grok Imagine | 50.1 [49.0, 51.1] | 49.2 [48.3, 50.1] |
| Veo 3.1 | 49.6 [48.9, 50.4] | 57.0 [56.1, 57.8] |
| Hunyuan-1.5 | 30.0 [28.8, 31.2] | 30.7 [29.5, 31.8] |
| LTX-2.3 | 32.1 [31.0, 33.2] | 39.8 [38.6, 40.9] |
| Mitty-EPIC | 61.5 [60.5, 62.5] | 56.1 [55.2, 57.0] |
| SkyReels-V3 | 31.5 [30.5, 32.6] | 34.6 [33.6, 35.6] |
| LongCat | 31.0 [29.9, 32.1] | 32.0 [30.9, 33.2] |
| Wan2.2 | 32.4 [31.4, 33.3] | 33.7 [32.7, 34.7] |
| Comparison | Gripper Δ Core [95% CI] | Hand Δ Core [95% CI] |
|---|---|---|
| Seedance − Wan2.7 | +0.79 [+0.08, +1.48] | +1.52 [+0.92, +2.16] |
| Seedance − Kling | +2.82 [+2.04, +3.61] | +2.89 [+2.15, +3.73] |
| Wan2.7 − Kling | +2.04 [+1.40, +2.66] | +1.37 [+0.66, +2.16] |
| Score | M1 | M2 | M3 | M4 | Transfer Subscore |
|---|---|---|---|---|---|
| Pearson r | 0.791 | 0.818 | 0.880 | 0.877 | 0.930 |
| Judge Pair | M1 | M2 | M3 | M4 |
|---|---|---|---|---|
| Gemini / Qwen | 0.695 | 0.691 | 0.741 | 0.874 |
| Gemini / GPT | 0.582 | 0.636 | 0.726 | 0.891 |
| Qwen / GPT | 0.656 | 0.690 | 0.809 | 0.892 |
| Model | Interface | Source input |
|---|---|---|
| Seedance 2.0 | Video | Full clip |
| Wan2.7 | Video | Full clip |
| Kling-V3 | Video | Full clip |
| Veo 3.1 | Frame | First + last |
| Grok Imagine Video | Frame | Up to 7 frames |
| HunyuanVideo 1.5-I2V | Frame | Ordered frame(s) |
| LTX-2.3 | Frame | Ordered frame(s) |
| Mitty-EPIC14B | Frame | Ordered frame(s) |
| SkyReels-V3-R2V | Frame | Ordered frame(s) |
| LongCat | Frame | Ordered frame(s) |
| Wan2.2 | Frame | First frame |
| Model | Decoded Size | Frames @ fps |
|---|---|---|
| Seedance 2.0 | 1280×720 | 121 @ 24 |
| Wan2.7 | 1280×720 | 150 @ 30 |
| Kling-V3 | 1280×720 | 121 @ 24 |
| Veo 3.1 | 1280×720 | 144 @ 24 |
| Grok Imagine Video | 1280×720 | 121 @ 24 |
| HunyuanVideo 1.5-I2V | 1280×720 | 121 @ 24 |
| LTX-2.3 | 1280×720 | 120 @ 24 |
| Mitty-EPIC14B | 1280×720 | 37 @ 8 |
| SkyReels-V3-R2V | 1280×720 | 121 @ 24 |
| LongCat | 832×480 | 80 @ 15 |
| Wan2.2 | 720P area | 121 @ 24 |
| Parallel-Jaw Gripper | Dexterous Hand | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Task | Model | Goal Completion | Action Completion | Contact Transfer | Embod. Correct. | Video Quality | Core | Goal Completion | Action Completion | Contact Transfer | Embod. Correct. | Video Quality | Core |
| Kling-V3 | 0.736 | 0.837 | 0.782 | 0.735 | 0.803 | 77.1 | 0.738 | 0.837 | 0.847 | 0.886 | 0.807 | 83.7 | |
| Seedance 2.0 | 0.762 | 0.857 | 0.836 | 0.824 | 0.798 | 82.1 | 0.775 | 0.853 | 0.880 | 0.908 | 0.802 | 86.1 | |
| Wan2.7 | 0.744 | 0.843 | 0.824 | 0.776 | 0.793 | 79.7 | 0.777 | 0.845 | 0.890 | 0.913 | 0.799 | 86.4 | |
| Grok Imagine | 0.608 | 0.704 | 0.555 | 0.430 | 0.806 | 57.3 | 0.656 | 0.742 | 0.561 | 0.368 | 0.811 | 57.0 | |
| Veo 3.1 | 0.751 | 0.791 | 0.574 | 0.197 | 0.790 | 54.2 | 0.780 | 0.830 | 0.653 | 0.305 | 0.800 | 60.9 | |
| Hunyuan-1.5 | 0.520 | 0.446 | 0.188 | 0.000 | 0.812 | 28.2 | 0.483 | 0.429 | 0.209 | 0.027 | 0.814 | 28.9 | |
| LTX-2.3 | 0.424 | 0.550 | 0.344 | 0.018 | 0.784 | 33.3 | 0.455 | 0.558 | 0.422 | 0.170 | 0.791 | 40.9 | |
| Mitty-EPIC | 0.522 | 0.591 | 0.540 | 0.575 | 0.733 | 57.5 | 0.528 | 0.640 | 0.591 | 0.402 | 0.733 | 54.6 | |
| SkyReels-V3 | 0.408 | 0.527 | 0.267 | 0.000 | 0.792 | 29.9 | 0.376 | 0.489 | 0.333 | 0.061 | 0.795 | 32.8 | |
| LongCat | 0.412 | 0.502 | 0.208 | 0.000 | 0.799 | 27.9 | 0.394 | 0.476 | 0.306 | 0.032 | 0.798 | 31.2 | |
| F1 Rigid | Wan2.2 | 0.377 | 0.547 | 0.229 | 0.000 | 0.776 | 28.5 | 0.422 | 0.553 | 0.249 | 0.000 | 0.774 | 29.8 |
| Kling-V3 | 0.641 | 0.792 | 0.686 | 0.682 | 0.802 | 70.6 | 0.666 | 0.796 | 0.770 | 0.885 | 0.804 | 79.6 | |
| Seedance 2.0 | 0.760 | 0.847 | 0.737 | 0.742 | 0.797 | 76.4 | 0.735 | 0.854 | 0.828 | 0.927 | 0.801 | 84.5 | |
| Wan2.7 | 0.695 | 0.783 | 0.704 | 0.771 | 0.794 | 74.4 | 0.676 | 0.795 | 0.778 | 0.913 | 0.799 | 80.8 | |
| Grok Imagine | 0.646 | 0.729 | 0.491 | 0.527 | 0.791 | 59.1 | 0.683 | 0.770 | 0.573 | 0.344 | 0.807 | 57.4 | |
| Veo 3.1 | 0.670 | 0.754 | 0.508 | 0.049 | 0.785 | 45.9 | 0.691 | 0.824 | 0.620 | 0.151 | 0.793 | 53.8 | |
| Hunyuan-1.5 | 0.433 | 0.517 | 0.178 | 0.014 | 0.807 | 28.1 | 0.399 | 0.481 | 0.202 | 0.123 | 0.809 | 31.0 | |
| LTX-2.3 | 0.478 | 0.514 | 0.278 | 0.000 | 0.774 | 31.0 | 0.523 | 0.613 | 0.328 | 0.157 | 0.786 | 39.4 | |
| Mitty-EPIC | 0.515 | 0.588 | 0.537 | 0.579 | 0.746 | 57.5 | 0.507 | 0.609 | 0.540 | 0.393 | 0.746 | 52.2 | |
| SkyReels-V3 | 0.454 | 0.552 | 0.223 | 0.000 | 0.790 | 29.7 | 0.425 | 0.559 | 0.303 | 0.011 | 0.794 | 32.1 | |
| LongCat | 0.454 | 0.579 | 0.243 | 0.000 | 0.792 | 30.7 | 0.403 | 0.577 | 0.275 | 0.011 | 0.794 | 31.2 | |
| F2 Mechanism | Wan2.2 | 0.427 | 0.607 | 0.234 | 0.000 | 0.763 | 30.2 | 0.526 | 0.672 | 0.264 | 0.000 | 0.762 | 33.5 |
| Kling-V3 | 0.683 | 0.748 | 0.748 | 0.735 | 0.794 | 73.9 | 0.680 | 0.739 | 0.796 | 0.907 | 0.799 | 80.4 | |
| Seedance 2.0 | 0.722 | 0.792 | 0.834 | 0.839 | 0.789 | 80.8 | 0.737 | 0.832 | 0.867 | 0.907 | 0.793 | 84.7 | |
| Wan2.7 | 0.623 | 0.727 | 0.742 | 0.785 | 0.789 | 73.9 | 0.660 | 0.762 | 0.813 | 0.912 | 0.792 | 81.0 | |
| Grok Imagine | 0.599 | 0.688 | 0.422 | 0.320 | 0.788 | 49.4 | 0.603 | 0.650 | 0.398 | 0.299 | 0.807 | 47.8 | |
| Veo 3.1 | 0.687 | 0.764 | 0.504 | 0.072 | 0.780 | 46.9 | 0.619 | 0.770 | 0.597 | 0.240 | 0.791 | 53.8 | |
| Hunyuan-1.5 | 0.462 | 0.448 | 0.214 | 0.000 | 0.803 | 28.1 | 0.409 | 0.469 | 0.193 | 0.022 | 0.805 | 27.7 | |
| LTX-2.3 | 0.426 | 0.470 | 0.276 | 0.046 | 0.773 | 30.8 | 0.510 | 0.516 | 0.343 | 0.195 | 0.779 | 39.3 |
| Metric | Evidence | Computation |
|---|---|---|
| M1: Goal | 25 generated frames and final-state predicates. | Weighted mean of normalized 0–4 predicate scores. |
| M2: Action | 25 generated frames and required events. | Weighted mean of normalized event-completion scores. |
| M3: Contact | 25 source frames, 25 generated frames, and the contact specification. | Mean of applicable contact dimensions; zero when source grounding fails. |
| M4: Embodiment | 25 generated frames and the target specification. | Weighted mean of five embodiment dimensions; zero on a hard failure. |
| M5: Quality | Decoded generated video. | Mean of MUSIQ, CLIP aesthetic, temporal-stability, and AMT-S scores. |
Findings
- The three video-conditioned models (Seedance 2.0, Wan2.7, Kling-V3) ranked at the top, with Seedance 2.0 scoring H2RCore 77.3 for the gripper and 84.6 for the dexterous hand, followed by Wan2.7 (76.5/83.1) and Kling-V3 (74.5/81.7).
- Video quality (M5) varied only slightly across models (0.73-0.81), but H2RCore ranged widely from 30.0 to 84.6, with a weak rank correlation between them (Spearman rho=0.14).
- Human raters and MLLM judges agreed strongly, with within-scene Spearman rho=0.883 and aggregate Pearson r=0.930 between human and automated transfer subscores.
- Nine of 11 models scored higher with the dexterous hand than the gripper, averaging a 3.3-point H2RCore gain, with the effect strongest for the three video-conditioned models (gains of 7.3, 6.6, and 7.2 points for Seedance, Wan2.7, and Kling-V3 respectively).
- Adding a target-robot reference image raised Wan2.7's H2RCore from 76.5 to 83.1 for the gripper, but lowered Kling-V3's by 13.5 points and Seedance's by 8.8 points, showing the effect was model-dependent rather than uniformly beneficial.
Where it can be used
- Teams building pipelines that convert human demonstration videos into robot training data can use this benchmark's five dimensions as a checklist beyond just visual quality, specifically checking contact and embodiment correctness.
- Anyone evaluating whether to adopt a video generation model for robot data augmentation can reference these five scoring axes (goal, action, contact, embodiment, quality) as evaluation criteria.
- When choosing between gripper and dexterous-hand target embodiments for such conversion tasks, this benchmark suggests the hand embodiment, being closer to the human hand shape, may transfer contact more reliably.
Limits and open work
- The benchmark evaluates only visible evidence in the generated video, not physical executability or downstream robot policy performance.
- It covers only 120 short EgoDex clips and two robot embodiments, capturing just part of the variation found in real manipulation settings.
- Models were evaluated through different native input interfaces (full video versus a few sampled frames), so model capability and interface differences are entangled in the comparison.
- MLLM judge scores can still be uncertain under occlusion, subtle contact, or severe generation artifacts, and extending the benchmark to longer demonstrations or richer 3D evidence is left for future work.
Why it matters
As people increasingly consider using video generation AI to convert abundant human demo videos into robot training data, this study shows with concrete numbers that current models often produce nice-looking videos that fail to actually transfer contact or the correct robot body type. Anyone planning to use such generated videos as robot training data needs to check task transfer and embodiment correctness separately, since visual quality alone does not indicate success.
Terms in this paper
- world model · an AI model that predicts future situations or generates video of what happens next; here video generators are treated as predictors of robot behavior
- embodiment · the physical body form of a robot, such as a parallel-jaw gripper or a dexterous five-fingered hand, that interacts with objects
- H2RCore · the paper's combined 0-100 score that weights the five evaluation dimensions (M1-M5) together
- MLLM judge · a large multimodal language model used to score videos automatically; this study used Gemini, Qwen, and GPT as three independent judges
- functional contact · whether the robot actually touches the object in a way that serves the same purpose as the human's touch, even if the grasp or path differs
Original abstract (English)
Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human manipulation videos provide rich behavioral experiences, but transferring them across embodiments remains challenging due to differences between human hands and robotic end-effectors. Recent advances in video world models offer a promising pathway to synthesize robot-centric manipulation videos from human observations, while their cross-embodiment transfer capability remains largely unexplored. Therefore, we introduce H2R-Bench, a benchmark for evaluating cross-embodiment human-to-robot manipulation video generation, where models transform egocentric human demonstrations into robot manipulation videos under specified embodiments. Each benchmark instance contains a human demonstration video, target embodiment constraints, and source-grounded annotations covering task goals, action events, functional contacts, and object responses. H2R-Bench evaluates generated videos through five dimensions, including goal-state completion, action-event completion, functional contact transfer, embodiment correctness, and general video quality. We benchmark eleven state-of-the-art video generation models across six manipulation families and two robot embodiments. Our evaluation reveals that current video world models remain limited in human-to-robot manipulation transfer: even leading models often fail in embodiment consistency, functional interaction, and task execution. H2R-Bench provides a systematic diagnostic framework for evaluating whether video world models can bridge the human-to-robot embodiment gap and convert human manipulation observations into robot-centric training resources.
Read on arXivLatest papers
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?AI coding agents were tested on fixing real scientific software, and even the best one failed more than half the time
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM ServingMaking sparse attention fast enough and accurate enough for real LLM serving, not just papers
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM AgentsMaking customer-service AI agents follow the whole procedure, not just avoid one bad action
- EXIMO: VLM Guided Exploration of VLA PoliciesTeaching a robot new chores without human teleoperation, by letting a chatty AI supervise it
- EnvHarness: Awakening Static Worlds for Agent LearningInstead of building new training worlds from scratch, this work adds a plug-in layer that reshapes existing ones around each agent's actual weaknesses
- Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the ModelCompanies that rent AI instead of owning it can only do half of AI safety oversight
- PersonalBench: Measuring the Authorship Gap in LLM PersonalizationAI can be prompted to write 'like someone,' but its own voice never fully disappears
- Automated Summarization of Financial News Using Large Language Models and Retrieval-Augmented Generation: An Early Empirical Study (Fall 2023)Testing AI summaries of stock news, the simple approach beat the trendy retrieval-based one
Latest from METAL MEDIA
Figures: Dingyi Rong et al., arXiv:2608.13049, arxiv-nonexclusive