H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models
研究者第一次系统检验:AI把人类操作视频改成机器人操作视频后,动作是否真的成功转移了
机器人训练用的示范视频稀缺又昂贵,而人类操作物体的第一人称视频却很丰富,于是有人尝试用视频生成AI把人类示范视频转换成机器人示范视频。以往的评测只看生成视频是否看起来逼真,却没人检查任务、接触方式和机器人形态是否真的正确转移了。H2R-Bench用120段人类示范视频,分别转换成两种机器人形态,对11个最先进的视频生成模型打分,发现视频画质好看和任务真正转移成功几乎没有关系。
METAL MEDIA 解读图
H2R-Bench 评测流程
证据状态已报告实测结果
- 输入:人类示范视频来自EgoDex的120段第一人称人类操作物体视频,均匀分布在六个任务类别中
- 条件:目标机器人形态每段源视频分别被要求转换为平行夹爪版本和灵巧五指手版本,共构成240个测试案例
- 生成:11个视频AI模型Seedance 2.0、Wan2.7、Kling-V3等完整视频输入模型与HunyuanVideo等帧输入模型各自通过原生接口生成机器人视频
- 打分:M1至M5五项指标目标完成度、动作完成度、功能性接触、机器人形态正确性、视频画质,均由三个MLLM评审按0到4分打分
- 结果:H2RCore综合得分五项分数加权汇总为0到100分,其中接触和形态正确性各占30%的权重
他们做了什么
- 研究团队从EgoDex数据集中挑选了120段第一人称人类操作物体的视频,要求11个视频生成AI模型把每段视频分别转换成'平行夹爪机器人版'和'灵巧五指机器人手版',共构成240个测试案例。
- 评测分五个维度打分:目标状态是否完成(M1)、所需动作是否都出现(M2)、机器人与物体是否建立了功能性接触(M3)、生成的机器人形态是否符合要求(M4),以及整体视频画质(M5)。三个大型多模态AI评审(Gemini、Qwen、GPT)独立打分,并与人类评分者的结果进行了对照验证。
- 能接收完整源视频作为输入的三个模型Seedance 2.0、Wan2.7、Kling-V3排名靠前,其中Seedance 2.0综合得分H2RCore(0到100分)最高,夹爪版77.3分,灵巧手版84.6分。
- 视频画质分数(M5)在各模型间差异很小,只在0.73到0.81之间,但综合转移得分H2RCore却在30.0到84.6之间大幅波动,两者的排名相关性很弱(斯皮尔曼相关系数仅0.14),说明画面好看不代表任务真的转移成功了。
- HunyuanVideo 1.5-I2V的视频画质分数最高,但接触得分很低,常常让人类的手继续完成操作而没有换成机器人;Veo 3.1能生成看起来合理的互动,但经常生成错误的末端执行器类型。

| Benchmark | Settings | Evaluation | ||||||
|---|---|---|---|---|---|---|---|---|
| I2V | RV | H2R | VQ | Goal | Action | Cont. | Emb. | |
| VBench | × | × | × | ✓ | × | × | × | × |
| WorldModelBench | ✓ | △ | × | ✓ | △ | △ | × | × |
| RBench | ✓ | ✓ | × | ✓ | ✓ | ✓ | × | △ |
| RoboWM-Bench | ✓ | ✓ | × | △ | ✓ | ✓ | × | × |
| RoboTrustBench | ✓ | ✓ | × | ✓ | ✓ | ✓ | × | △ |
| H2R-Bench (ours) | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |

| Model | Parallel-Jaw Gripper | Dexterous Hand | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Component Metrics | Aggregate | Component Metrics | Aggregate | |||||||||
| Goal Comp. | Action Comp. | Contact Transfer | Embod. Correct. | Video Quality | H2R Core | Goal Comp. | Action Comp. | Contact Transfer | Embod. Correct. | Video Quality | H2R Core | |
| Video-conditioned generation | ||||||||||||
| Seedance 2.0 | 0.725 | 0.813 | 0.776 | 0.768 | 0.793 | 77.3 | 0.744 | 0.832 | 0.855 | 0.911 | 0.799 | 84.6 |
| Wan2.7 | 0.706 | 0.791 | 0.766 | 0.772 | 0.796 | 76.5 | 0.718 | 0.804 | 0.835 | 0.910 | 0.795 | 83.1 |
| Kling-V3 | 0.710 | 0.807 | 0.751 | 0.707 | 0.798 | 74.5 | 0.707 | 0.800 | 0.819 | 0.885 | 0.802 | 81.7 |
| Frame-conditioned generation | ||||||||||||
| Mitty-EPIC14B | 0.581 | 0.668 | 0.598 | 0.585 | 0.732 | 61.5 | 0.587 | 0.684 | 0.598 | 0.392 | 0.732 | 56.1 |
| Veo 3.1 | 0.725 | 0.797 | 0.533 | 0.100 | 0.783 | 49.6 | 0.715 | 0.816 | 0.642 | 0.227 | 0.793 | 57.0 |
| Grok Imagine Video | 0.661 | 0.729 | 0.443 | 0.268 | 0.792 | 50.1 | 0.678 | 0.729 | 0.469 | 0.198 | 0.804 | 49.2 |
| LTX-2.3 | 0.473 | 0.545 | 0.292 | 0.012 | 0.773 | 32.1 | 0.520 | 0.592 | 0.377 | 0.132 | 0.780 | 39.8 |
| SkyReels-V3-R2V | 0.448 | 0.610 | 0.256 | 0.004 | 0.787 | 31.5 | 0.441 | 0.607 | 0.341 | 0.026 | 0.789 | 34.6 |
| Wan2.2 | 0.492 | 0.639 | 0.258 | 0.000 | 0.769 | 32.4 | 0.512 | 0.653 | 0.286 | 0.000 | 0.766 | 33.7 |
| LongCat | 0.472 | 0.583 | 0.243 | 0.000 | 0.790 | 31.0 | 0.428 | 0.545 | 0.298 | 0.020 | 0.793 | 32.0 |
| HunyuanVideo 1.5-I2V | 0.535 | 0.549 | 0.184 | 0.005 | 0.806 | 30.0 | 0.499 | 0.555 | 0.185 | 0.041 | 0.808 | 30.7 |

| Metric | Mean change (Hand − Gripper) | Hand higher (models) |
|---|---|---|
| Goal-State Completion (M1) | +0.002 | 6/11 |
| Action-Event Completion (M2) | +0.008 | 8/11 |
| Functional Contact Transfer (M3) | +0.055 | 11/11 |
| Embodiment Correctness (M4) | +0.047 | 8/11 |
| Video Quality (M5) | +0.004 | 8/11 |
| H2RCore (0–100) | +3.3 | 9/11 |
| Model | Ref. | Goal | Action | Contact | Embod. | Quality | Core |
|---|---|---|---|---|---|---|---|
| Parallel-Jaw Gripper | |||||||
| Kling-V3 | No | 0.710 | 0.807 | 0.751 | 0.707 | 0.798 | 74.5 |
| Yes | 0.673 | 0.760 | 0.577 | 0.479 | 0.779 | 61.0 | |
| Seedance 2.0 | No | 0.725 | 0.813 | 0.776 | 0.768 | 0.793 | 77.3 |
| Yes | 0.711 | 0.794 | 0.674 | 0.595 | 0.784 | 68.5 | |
| Wan2.7 | No | 0.706 | 0.791 | 0.766 | 0.772 | 0.796 | 76.5 |
| Yes | 0.714 | 0.811 | 0.871 | 0.875 | 0.781 | 83.1 | |
| Dexterous Hand | |||||||
| Kling-V3 | No | 0.707 | 0.800 | 0.819 | 0.885 | 0.802 | 81.7 |
| Yes | 0.629 | 0.734 | 0.772 | 0.728 | 0.799 | 73.4 | |
| Seedance 2.0 | No | 0.744 | 0.832 | 0.855 | 0.911 | 0.799 | 84.6 |
| Yes | 0.736 | 0.824 | 0.876 | 0.753 | 0.794 | 80.2 | |
| Wan2.7 | No | 0.718 | 0.804 | 0.835 | 0.910 | 0.795 | 83.1 |
| Yes | 0.712 | 0.819 | 0.905 | 0.873 | 0.789 | 84.2 |

| Element | Embodied-manipulation concern | H2R-Bench operationalization | Design origin |
|---|---|---|---|
| Task families: task-defining physical state changes | |||
| F1: Rigid rearrangement | Object pose, support, or containment | Place or transport a rigid object into the demonstrated spatial relation. | Rigid transport and placement tasks. |
| F2: Mechanism actuation | State of an articulated mechanism | Open, close, toggle, press, or rotate a task-relevant mechanism. | Interaction with articulated objects. |
| F3: Insertion and assembly | Connection, fit, or attachment relation | Establish or remove a constrained connection between entities. | Precision alignment and constrained contact. |
| F4: Deformable configuration | Non-rigid shape or configuration | Produce the demonstrated fold, bend, compression, or shape change. | Deformable-object manipulation. |
| F5: Bulk-material transfer | Distribution or containment of material | Pour, scoop, transfer, or mix material between regions or containers. | Many-particle and material-flow manipulation. |
| F6: Surface/material transformation | Local surface condition or material integrity | Produce a visible local change through wiping, spreading, cutting, peeling, or related interaction. | Tool-mediated local transformation. |
| Evaluation dimensions: evidence required for valid H2R generation | |||
| M1: Goal-State Completion | Was the demonstrated task state reached? | Score weighted predicates over the visible final state. | Task-success and outcome evaluation in robotic video benchmarks. |
| M2: Action-Event Completion | Were the required manipulation events shown? | Score completion of source-derived action events. | Action-completeness evaluation in robotic video benchmarks. |
| M3: Functional Contact Transfer | Did robot contact support the source-consistent object response? | Evaluate contact region, establishment, mode, temporal object response, and embodiment-compatible strategy. | Contact-mediated manipulation and H2R interaction grounding. |
| M4: Embodiment Correctness | Did the requested robot perform the manipulation? | Evaluate robot presence, human absence, embodiment category, end-effector subtype, and temporal structure. | Robot-structure and embodiment compliance. |
| M5: Video Quality | Is the generated video visually and temporally well formed? | Measure imaging quality, aesthetic quality, temporal stability, and motion smoothness. | Task-agnostic video-generation quality. |
| Stage | Output | Quality control |
|---|---|---|
| Source curation | Curated clips across six families | Identifiable manipulated entities, visible task-relevant state change, and sufficient interaction evidence. |
| Clip grounding | 32 ordered source frames per clip | First and last frames are retained; annotations identify initial state, contact, core transition, and final state when visible. |
| Structured annotation | Goals, events, milestones, interaction, and embodiment strategies | Strict JSON schema with canonical event identifiers and a fixed six-family taxonomy. |
| Completion and validation | Schema-complete case records | Only missing fields are completed; existing non-empty fields are preserved; schema and type mismatches are rejected. |
| Manual verification | All benchmark annotations | Every record is checked against its source video and corrected before evaluation. |
| Uncertainty accounting | Validity and ambiguity fields | Every record has a validity decision; notes document residual uncertainty without changing the official evaluation set. |
| Target | Input | M1 | M2 | M3 | M4 | M5 | Core |
|---|---|---|---|---|---|---|---|
| Parallel-Jaw Gripper | Video | 0.737 | 0.850 | 0.709 | 0.598 | 0.793 | 70.9 |
| 9 images | 0.744 | 0.848 | 0.284 | 0.104 | 0.783 | 43.4 | |
| Dexterous Hand | Video | 0.768 | 0.857 | 0.861 | 0.898 | 0.799 | 85.1 |
| 9 images | 0.753 | 0.856 | 0.313 | 0.079 | 0.782 | 43.7 |
| Model | Parallel-Jaw Gripper Core [95% CI] | Dexterous Hand Core [95% CI] |
|---|---|---|
| Kling-V3 | 74.5 [73.6, 75.3] | 81.7 [80.8, 82.6] |
| Seedance 2.0 | 77.3 [76.5, 78.1] | 84.6 [84.0, 85.3] |
| Wan2.7 | 76.5 [75.7, 77.4] | 83.1 [82.3, 83.9] |
| Grok Imagine | 50.1 [49.0, 51.1] | 49.2 [48.3, 50.1] |
| Veo 3.1 | 49.6 [48.9, 50.4] | 57.0 [56.1, 57.8] |
| Hunyuan-1.5 | 30.0 [28.8, 31.2] | 30.7 [29.5, 31.8] |
| LTX-2.3 | 32.1 [31.0, 33.2] | 39.8 [38.6, 40.9] |
| Mitty-EPIC | 61.5 [60.5, 62.5] | 56.1 [55.2, 57.0] |
| SkyReels-V3 | 31.5 [30.5, 32.6] | 34.6 [33.6, 35.6] |
| LongCat | 31.0 [29.9, 32.1] | 32.0 [30.9, 33.2] |
| Wan2.2 | 32.4 [31.4, 33.3] | 33.7 [32.7, 34.7] |
| Comparison | Gripper Δ Core [95% CI] | Hand Δ Core [95% CI] |
|---|---|---|
| Seedance − Wan2.7 | +0.79 [+0.08, +1.48] | +1.52 [+0.92, +2.16] |
| Seedance − Kling | +2.82 [+2.04, +3.61] | +2.89 [+2.15, +3.73] |
| Wan2.7 − Kling | +2.04 [+1.40, +2.66] | +1.37 [+0.66, +2.16] |
| Score | M1 | M2 | M3 | M4 | Transfer Subscore |
|---|---|---|---|---|---|
| Pearson r | 0.791 | 0.818 | 0.880 | 0.877 | 0.930 |
| Judge Pair | M1 | M2 | M3 | M4 |
|---|---|---|---|---|
| Gemini / Qwen | 0.695 | 0.691 | 0.741 | 0.874 |
| Gemini / GPT | 0.582 | 0.636 | 0.726 | 0.891 |
| Qwen / GPT | 0.656 | 0.690 | 0.809 | 0.892 |
| Model | Interface | Source input |
|---|---|---|
| Seedance 2.0 | Video | Full clip |
| Wan2.7 | Video | Full clip |
| Kling-V3 | Video | Full clip |
| Veo 3.1 | Frame | First + last |
| Grok Imagine Video | Frame | Up to 7 frames |
| HunyuanVideo 1.5-I2V | Frame | Ordered frame(s) |
| LTX-2.3 | Frame | Ordered frame(s) |
| Mitty-EPIC14B | Frame | Ordered frame(s) |
| SkyReels-V3-R2V | Frame | Ordered frame(s) |
| LongCat | Frame | Ordered frame(s) |
| Wan2.2 | Frame | First frame |
| Model | Decoded Size | Frames @ fps |
|---|---|---|
| Seedance 2.0 | 1280×720 | 121 @ 24 |
| Wan2.7 | 1280×720 | 150 @ 30 |
| Kling-V3 | 1280×720 | 121 @ 24 |
| Veo 3.1 | 1280×720 | 144 @ 24 |
| Grok Imagine Video | 1280×720 | 121 @ 24 |
| HunyuanVideo 1.5-I2V | 1280×720 | 121 @ 24 |
| LTX-2.3 | 1280×720 | 120 @ 24 |
| Mitty-EPIC14B | 1280×720 | 37 @ 8 |
| SkyReels-V3-R2V | 1280×720 | 121 @ 24 |
| LongCat | 832×480 | 80 @ 15 |
| Wan2.2 | 720P area | 121 @ 24 |
| Parallel-Jaw Gripper | Dexterous Hand | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Task | Model | Goal Completion | Action Completion | Contact Transfer | Embod. Correct. | Video Quality | Core | Goal Completion | Action Completion | Contact Transfer | Embod. Correct. | Video Quality | Core |
| Kling-V3 | 0.736 | 0.837 | 0.782 | 0.735 | 0.803 | 77.1 | 0.738 | 0.837 | 0.847 | 0.886 | 0.807 | 83.7 | |
| Seedance 2.0 | 0.762 | 0.857 | 0.836 | 0.824 | 0.798 | 82.1 | 0.775 | 0.853 | 0.880 | 0.908 | 0.802 | 86.1 | |
| Wan2.7 | 0.744 | 0.843 | 0.824 | 0.776 | 0.793 | 79.7 | 0.777 | 0.845 | 0.890 | 0.913 | 0.799 | 86.4 | |
| Grok Imagine | 0.608 | 0.704 | 0.555 | 0.430 | 0.806 | 57.3 | 0.656 | 0.742 | 0.561 | 0.368 | 0.811 | 57.0 | |
| Veo 3.1 | 0.751 | 0.791 | 0.574 | 0.197 | 0.790 | 54.2 | 0.780 | 0.830 | 0.653 | 0.305 | 0.800 | 60.9 | |
| Hunyuan-1.5 | 0.520 | 0.446 | 0.188 | 0.000 | 0.812 | 28.2 | 0.483 | 0.429 | 0.209 | 0.027 | 0.814 | 28.9 | |
| LTX-2.3 | 0.424 | 0.550 | 0.344 | 0.018 | 0.784 | 33.3 | 0.455 | 0.558 | 0.422 | 0.170 | 0.791 | 40.9 | |
| Mitty-EPIC | 0.522 | 0.591 | 0.540 | 0.575 | 0.733 | 57.5 | 0.528 | 0.640 | 0.591 | 0.402 | 0.733 | 54.6 | |
| SkyReels-V3 | 0.408 | 0.527 | 0.267 | 0.000 | 0.792 | 29.9 | 0.376 | 0.489 | 0.333 | 0.061 | 0.795 | 32.8 | |
| LongCat | 0.412 | 0.502 | 0.208 | 0.000 | 0.799 | 27.9 | 0.394 | 0.476 | 0.306 | 0.032 | 0.798 | 31.2 | |
| F1 Rigid | Wan2.2 | 0.377 | 0.547 | 0.229 | 0.000 | 0.776 | 28.5 | 0.422 | 0.553 | 0.249 | 0.000 | 0.774 | 29.8 |
| Kling-V3 | 0.641 | 0.792 | 0.686 | 0.682 | 0.802 | 70.6 | 0.666 | 0.796 | 0.770 | 0.885 | 0.804 | 79.6 | |
| Seedance 2.0 | 0.760 | 0.847 | 0.737 | 0.742 | 0.797 | 76.4 | 0.735 | 0.854 | 0.828 | 0.927 | 0.801 | 84.5 | |
| Wan2.7 | 0.695 | 0.783 | 0.704 | 0.771 | 0.794 | 74.4 | 0.676 | 0.795 | 0.778 | 0.913 | 0.799 | 80.8 | |
| Grok Imagine | 0.646 | 0.729 | 0.491 | 0.527 | 0.791 | 59.1 | 0.683 | 0.770 | 0.573 | 0.344 | 0.807 | 57.4 | |
| Veo 3.1 | 0.670 | 0.754 | 0.508 | 0.049 | 0.785 | 45.9 | 0.691 | 0.824 | 0.620 | 0.151 | 0.793 | 53.8 | |
| Hunyuan-1.5 | 0.433 | 0.517 | 0.178 | 0.014 | 0.807 | 28.1 | 0.399 | 0.481 | 0.202 | 0.123 | 0.809 | 31.0 | |
| LTX-2.3 | 0.478 | 0.514 | 0.278 | 0.000 | 0.774 | 31.0 | 0.523 | 0.613 | 0.328 | 0.157 | 0.786 | 39.4 | |
| Mitty-EPIC | 0.515 | 0.588 | 0.537 | 0.579 | 0.746 | 57.5 | 0.507 | 0.609 | 0.540 | 0.393 | 0.746 | 52.2 | |
| SkyReels-V3 | 0.454 | 0.552 | 0.223 | 0.000 | 0.790 | 29.7 | 0.425 | 0.559 | 0.303 | 0.011 | 0.794 | 32.1 | |
| LongCat | 0.454 | 0.579 | 0.243 | 0.000 | 0.792 | 30.7 | 0.403 | 0.577 | 0.275 | 0.011 | 0.794 | 31.2 | |
| F2 Mechanism | Wan2.2 | 0.427 | 0.607 | 0.234 | 0.000 | 0.763 | 30.2 | 0.526 | 0.672 | 0.264 | 0.000 | 0.762 | 33.5 |
| Kling-V3 | 0.683 | 0.748 | 0.748 | 0.735 | 0.794 | 73.9 | 0.680 | 0.739 | 0.796 | 0.907 | 0.799 | 80.4 | |
| Seedance 2.0 | 0.722 | 0.792 | 0.834 | 0.839 | 0.789 | 80.8 | 0.737 | 0.832 | 0.867 | 0.907 | 0.793 | 84.7 | |
| Wan2.7 | 0.623 | 0.727 | 0.742 | 0.785 | 0.789 | 73.9 | 0.660 | 0.762 | 0.813 | 0.912 | 0.792 | 81.0 | |
| Grok Imagine | 0.599 | 0.688 | 0.422 | 0.320 | 0.788 | 49.4 | 0.603 | 0.650 | 0.398 | 0.299 | 0.807 | 47.8 | |
| Veo 3.1 | 0.687 | 0.764 | 0.504 | 0.072 | 0.780 | 46.9 | 0.619 | 0.770 | 0.597 | 0.240 | 0.791 | 53.8 | |
| Hunyuan-1.5 | 0.462 | 0.448 | 0.214 | 0.000 | 0.803 | 28.1 | 0.409 | 0.469 | 0.193 | 0.022 | 0.805 | 27.7 | |
| LTX-2.3 | 0.426 | 0.470 | 0.276 | 0.046 | 0.773 | 30.8 | 0.510 | 0.516 | 0.343 | 0.195 | 0.779 | 39.3 |
| Metric | Evidence | Computation |
|---|---|---|
| M1: Goal | 25 generated frames and final-state predicates. | Weighted mean of normalized 0–4 predicate scores. |
| M2: Action | 25 generated frames and required events. | Weighted mean of normalized event-completion scores. |
| M3: Contact | 25 source frames, 25 generated frames, and the contact specification. | Mean of applicable contact dimensions; zero when source grounding fails. |
| M4: Embodiment | 25 generated frames and the target specification. | Weighted mean of five embodiment dimensions; zero on a hard failure. |
| M5: Quality | Decoded generated video. | Mean of MUSIQ, CLIP aesthetic, temporal-stability, and AMT-S scores. |
研究结果
- 三个可接收完整源视频的模型(Seedance 2.0、Wan2.7、Kling-V3)排名靠前,Seedance 2.0的H2RCore得分夹爪版77.3、灵巧手版84.6,Wan2.7为76.5/83.1,Kling-V3为74.5/81.7。
- 视频画质分数(M5)在各模型间只有0.73到0.81的小幅差异,但H2RCore综合得分却在30.0到84.6之间大幅波动,两者排名相关性很弱(斯皮尔曼相关系数为0.14)。
- 人类评分者与AI评审的评分高度一致,场景内排名的斯皮尔曼相关系数为0.883,人类与自动化转移子分数的整体皮尔逊相关系数为0.930。
- 11个模型中有9个在灵巧手条件下得分高于夹爪条件,平均H2RCore提高3.3分,其中三个可接收完整视频的模型提升最明显,分别为Seedance提升7.3分、Wan2.7提升6.6分、Kling-V3提升7.2分。
- 增加目标机器人参考图像后,Wan2.7的H2RCore从76.5提升到83.1,但Kling-V3下降了13.5分、Seedance下降了8.8分,说明这种参考图像的效果因模型而异,并非普遍有益。
可应用场景
- 构建将人类示范视频转换为机器人训练数据的流水线时,可以参考本文的五个评测维度作为检查清单,而不只看视频画质,特别要单独核查接触与机器人形态是否正确。
- 在评估是否采用某个视频生成模型作为机器人数据增强工具前,可以参考这五个评分维度(目标、动作、接触、形态、画质)作为自身评测标准。
- 在为此类转换任务选择夹爪还是灵巧手作为目标机器人形态时,本文结果提示形态更接近人手的灵巧手可能在接触转移上更有优势。
局限与待验证事项
- 该基准只评测生成视频中可见的证据,并不检验机器人动作在物理上是否真的可执行,也不检验下游机器人策略的实际表现。
- 研究只涵盖来自EgoDex的120段短视频和两种机器人形态,只能反映真实操作场景多样性的一部分。
- 各模型接收源视频的方式不同(完整视频或几张采样帧),因此模型能力差异和输入接口差异混杂在了比较结果中。
- 在遮挡、接触细微或生成明显失真的情况下,AI评审的打分仍可能存在不确定性,将基准扩展到更长示范视频或引入更丰富的三维证据留待未来工作。
为什么重要
越来越多人考虑用视频生成AI把丰富的人类操作视频转换成机器人训练数据,这项研究用具体数字表明,目前的模型往往只生成好看的视频,却没能真正让机器人建立接触或呈现要求的机器人形态。想要把这类生成视频用作机器人训练资源的人,必须单独检查接触和形态是否正确,而不能只看画质或表面上的合理性。
本文术语
- 世界模型(world model) · 能预测未来情境或生成对应视频的AI模型,这里把视频生成模型当作预测机器人未来动作的工具
- 机器人形态(embodiment) · 机器人操作物体所用的身体结构,例如平行夹爪或灵巧五指机器人手
- H2RCore · 本文将五个评测维度(M1到M5)加权汇总得出的0到100分综合得分
- MLLM评审 · 用来给视频打分的大型多模态语言模型,本文使用了Gemini、Qwen、GPT三个评审
- 功能性接触(functional contact) · 指机器人是否真的以能达成操作目的的方式接触了物体,即使抓握方式或路径与人类不同也算成功
论文原文摘要(英文)
Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human manipulation videos provide rich behavioral experiences, but transferring them across embodiments remains challenging due to differences between human hands and robotic end-effectors. Recent advances in video world models offer a promising pathway to synthesize robot-centric manipulation videos from human observations, while their cross-embodiment transfer capability remains largely unexplored. Therefore, we introduce H2R-Bench, a benchmark for evaluating cross-embodiment human-to-robot manipulation video generation, where models transform egocentric human demonstrations into robot manipulation videos under specified embodiments. Each benchmark instance contains a human demonstration video, target embodiment constraints, and source-grounded annotations covering task goals, action events, functional contacts, and object responses. H2R-Bench evaluates generated videos through five dimensions, including goal-state completion, action-event completion, functional contact transfer, embodiment correctness, and general video quality. We benchmark eleven state-of-the-art video generation models across six manipulation families and two robot embodiments. Our evaluation reveals that current video world models remain limited in human-to-robot manipulation transfer: even leading models often fail in embodiment consistency, functional interaction, and task execution. H2R-Bench provides a systematic diagnostic framework for evaluating whether video world models can bridge the human-to-robot embodiment gap and convert human manipulation observations into robot-centric training resources.
在 arXiv 阅读最新论文
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?让AI编程助手去修复真实科学软件,连最强的那个也有一半以上任务没做对
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving把稀疏注意力从论文原型变成能真正上线服务的加速方案
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents让客服AI坐席不只是拦住一个危险动作,而是把整个流程走对
- EXIMO: VLM Guided Exploration of VLA Policies不用人工遥控演示,让会说话的AI来教机械臂做新家务
- EnvHarness: Awakening Static Worlds for Agent Learning不重新搭建训练环境,而是给现有环境套一层可插拔组件,针对每个智能体的具体弱点重新塑形
- Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the Model租用AI而非拥有AI的机构,安全监管能力只剩一半
- Beyond Imitation: Filtering On-Policy Distillation by Reasoning ProgressAI模仿老师模型学习时,会误伤本来推理正确的步骤,新方法专门过滤掉这种误伤
- PersonalBench: Measuring the Authorship Gap in LLM Personalization让AI模仿某人的文风,结果发现它始终摆脱不了自己的腔调
METAL MEDIA 最新报道
图片来源: Dingyi Rong et al., arXiv:2608.13049, arxiv-nonexclusive