컴백부터 K-뷰티까지 — K-컬쳐의 모든 것을 메일로 받아보세요메일로 받아보기

METAL MEDIA

사람이 물건을 만지는 영상을 보고 AI가 만든 '로봇 버전 영상'이 실제로 쓸만한지 처음으로 채점해봤다

arXiv:2608.130492026-08-12

H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models

사람이 물건을 만지는 영상을 보고 AI가 만든 '로봇 버전 영상'이 실제로 쓸만한지 처음으로 채점해봤다

로봇을 학습시킬 영상은 부족한데 사람이 물건 다루는 영상은 넘쳐나서, AI 비디오 생성 모델로 사람 영상을 로봇 영상으로 바꾸려는 시도가 늘고 있다. 문제는 지금까지는 그렇게 만들어진 영상이 '그럴듯해 보이는지'만 평가했고, 실제로 사람이 하던 동작과 접촉이 로봇에게 제대로 옮겨졌는지는 아무도 확인하지 않았다는 점이다. H2R-Bench는 사람 시범 영상 120개를 로봇용 그리퍼와 다섯손가락 손 두 종류로 변환시켜 11개 최신 영상 생성 모델을 5가지 기준으로 채점했고, 화면이 예뻐 보이는 것과 실제 작업 전이 성공은 거의 무관하다는 것을 보여줬다.

METAL MEDIA 해설 도표

H2R-Bench 평가 흐름

증거 상태측정 결과가 보고됨

  1. 입력: 사람 시범 영상EgoDex에서 고른 120개 1인칭 사람 물건 조작 영상, 6개 작업군으로 균등 분배
  2. 조건: 목표 로봇 형태 지정각 영상을 평행 그리퍼용, 다섯손가락 로봇 손용 두 가지 목표로 각각 변환 요청, 총 240개 케이스
  3. 생성: 11개 영상 AI 모델Seedance 2.0, Wan2.7, Kling-V3 등 영상 전체 입력 모델과 HunyuanVideo 등 프레임 입력 모델이 각자 인터페이스로 로봇 영상 생성
  4. 채점: M1~M5 다섯 기준목표 완성, 동작 완성, 기능적 접촉, 로봇 형태 정확성, 영상 품질을 3개의 AI 판정관이 0~4점으로 평가
  5. 결과: H2RCore 종합점수다섯 점수를 가중합해 0~100점으로 산출, 접촉과 형태 정확성에 각 30%씩 가중을 둠
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.

무엇을 했나

  1. 연구팀은 사람이 1인칭 시점으로 물건을 다루는 영상(EgoDex 데이터셋에서 120개 클립)을 골라, 각 영상을 '평행 그리퍼용 로봇 영상'과 '다섯손가락 로봇 손 영상' 두 가지로 바꾸라고 11개 영상 생성 AI에게 시켰다.
  2. 평가는 다섯 항목으로 나눠 했다: 목표 상태 완성(M1), 필요한 동작이 다 나왔는지(M2), 로봇이 물건과 실제로 기능적으로 접촉했는지(M3), 요청한 로봇 형태(그리퍼/손)가 제대로 나왔는지(M4), 일반적인 영상 품질(M5). 세 개의 대형 AI 판정관(Gemini, Qwen, GPT)이 채점하고 사람 평가자와 비교해 신뢰도를 검증했다.
  3. Seedance 2.0, Wan2.7, Kling-V3처럼 원본 영상 전체를 입력받는 모델들이 상위권을 차지했고, Seedance 2.0이 종합점수(H2RCore, 0~100점)에서 그리퍼 77.3점, 다섯손가락 손 84.6점으로 1위였다.
  4. 영상 품질 점수(M5)는 모델 간 차이가 0.73~0.81로 거의 비슷했지만, 실제 작업 전이 점수(H2RCore)는 30.0~84.6점까지 크게 벌어졌고 둘 사이 상관관계는 매우 약했다(순위 상관 0.14). 즉 영상이 화려해 보인다고 로봇 작업이 제대로 전이된 건 아니었다.
  5. 화면이 좋아 보이는 HunyuanVideo 1.5-I2V는 영상 품질 점수가 가장 높았지만 로봇 접촉 점수는 낮았고, 심지어 사람 손이 계속 작업을 하는 영상을 만들어 로봇으로 바꾸지 못한 경우가 많았다. Veo 3.1은 그럴듯한 상호작용을 만들었지만 요청한 로봇 형태가 틀린 경우가 잦았다.
Figure 1: Comparison between existing evaluation and H2R-Bench. Given a human demonstration and a target embodiment instruction, video world models generate robot manipulation videos. Existing video benchmarks such as WorldModelBench (Li et al. 2026a) assess overall video plausibility, rates both videos highly, whereas our H2R-Bench diagnoses transfer through goal, action, contact, and embodiment.
Figure 1: Comparison between existing evaluation and H2R-Bench. Given a human demonstration and a target embodiment instruction, video world models generate robot manipulation videos. Existing video benchmarks such as WorldModelBench (Li et al. 2026a) assess overall video plausibility, rates both videos highly, whereas our H2R-Bench diagnoses transfer through goal, action, contact, and embodiment.
Table 1: Comparison of H2R-Bench and existing video-generation benchmarks across evaluation capabilities. “I2V”, “RV”, and “H2R” denote image-to-video generation, robot-video evaluation, and human-to-robot transfer, respectively. Evaluation dimensions include visual quality, goal completion, action completion, functional contact transfer, and embodiment consistency. ✓, △, and × indicate full, partial, and no support, respectively.
BenchmarkSettingsEvaluation
I2VRVH2RVQGoalActionCont.Emb.
VBench×××××××
WorldModelBench×××
RBench××
RoboWM-Bench×××
RoboTrustBench××
H2R-Bench (ours)
Figure 2: Overview of H2R-Bench. The benchmark curates egocentric human manipulation demonstrations, conditions video generators on each model’s supported source interface and target embodiment, and evaluates the resulting robot videos with transfer-aware metrics. The lower panels summarize task coverage and model capability profiles across both target embodiments.
Figure 2: Overview of H2R-Bench. The benchmark curates egocentric human manipulation demonstrations, conditions video generators on each model’s supported source interface and target embodiment, and evaluates the resulting robot videos with transfer-aware metrics. The lower panels summarize task coverage and model capability profiles across both target embodiments.
Table 2: Main-evaluation results by target embodiment. M1–M5 measure goal completion, action completion, contact transfer, embodiment correctness, and Video Quality; H2RCore aggregates all five metrics on a 0–100 scale. Rows are ordered by the sum of the two H2RCore scores within each conditioning group. Bold and underlined entries indicate the best and second-best result in each column.
ModelParallel-Jaw GripperDexterous Hand
Component MetricsAggregateComponent MetricsAggregate
Goal Comp.Action Comp.Contact TransferEmbod. Correct.Video QualityH2R CoreGoal Comp.Action Comp.Contact TransferEmbod. Correct.Video QualityH2R Core
Video-conditioned generation
Seedance 2.00.7250.8130.7760.7680.79377.30.7440.8320.8550.9110.79984.6
Wan2.70.7060.7910.7660.7720.79676.50.7180.8040.8350.9100.79583.1
Kling-V30.7100.8070.7510.7070.79874.50.7070.8000.8190.8850.80281.7
Frame-conditioned generation
Mitty-EPIC14B0.5810.6680.5980.5850.73261.50.5870.6840.5980.3920.73256.1
Veo 3.10.7250.7970.5330.1000.78349.60.7150.8160.6420.2270.79357.0
Grok Imagine Video0.6610.7290.4430.2680.79250.10.6780.7290.4690.1980.80449.2
LTX-2.30.4730.5450.2920.0120.77332.10.5200.5920.3770.1320.78039.8
SkyReels-V3-R2V0.4480.6100.2560.0040.78731.50.4410.6070.3410.0260.78934.6
Wan2.20.4920.6390.2580.0000.76932.40.5120.6530.2860.0000.76633.7
LongCat0.4720.5830.2430.0000.79031.00.4280.5450.2980.0200.79332.0
HunyuanVideo 1.5-I2V0.5350.5490.1840.0050.80630.00.4990.5550.1850.0410.80830.7
Figure 3: Agreement between human and MLLM evaluators. Human raters and MLLM judges rank generated videos based on the transfer score aggregated from M1–M4. MLLM-based evaluation closely aligns with human judgments, with Spearman correlations above 0.8 across evaluators.
Figure 3: Agreement between human and MLLM evaluators. Human raters and MLLM judges rank generated videos based on the transfer score aggregated from M1–M4. MLLM-based evaluation closely aligns with human judgments, with Spearman correlations above 0.8 across evaluators.
Table 3: Effect of target embodiment across all 11 models. Mean change reports the average score difference between the Dexterous Hand and Parallel-Jaw Gripper. “Hand higher” reports the number of models with a positive difference.
MetricMean change (Hand − Gripper)Hand higher (models)
Goal-State Completion (M1)+0.0026/11
Action-Event Completion (M2)+0.0088/11
Functional Contact Transfer (M3)+0.05511/11
Embodiment Correctness (M4)+0.0478/11
Video Quality (M5)+0.0048/11
H2RCore (0–100)+3.39/11
Figure 4: VBench Video Quality versus H2RCore. Unlike VBench, H2RCore better differentiates the models.
Figure 4: VBench Video Quality versus H2RCore. Unlike VBench, H2RCore better differentiates the models.
Table 4: Effects of target-robot reference images for three video-conditioned models. “No” and “Yes” indicate generation without and with a target-robot reference image.
ModelRef.GoalActionContactEmbod.QualityCore
Parallel-Jaw Gripper
Kling-V3No0.7100.8070.7510.7070.79874.5
Yes0.6730.7600.5770.4790.77961.0
Seedance 2.0No0.7250.8130.7760.7680.79377.3
Yes0.7110.7940.6740.5950.78468.5
Wan2.7No0.7060.7910.7660.7720.79676.5
Yes0.7140.8110.8710.8750.78183.1
Dexterous Hand
Kling-V3No0.7070.8000.8190.8850.80281.7
Yes0.6290.7340.7720.7280.79973.4
Seedance 2.0No0.7440.8320.8550.9110.79984.6
Yes0.7360.8240.8760.7530.79480.2
Wan2.7No0.7180.8040.8350.9100.79583.1
Yes0.7120.8190.9050.8730.78984.2
Figure 5: Qualitative H2R transfer results with the shared source video and prompt. The unscored top row is the human source; the generated rows show representative stages of each output. Metric strips report per-video M1–M5 and H2RCore.
Figure 5: Qualitative H2R transfer results with the shared source video and prompt. The unscored top row is the human source; the generated rows show representative stages of each output. Metric strips report per-video M1–M5 and H2RCore.
Table S1: Design rationale for the H2R-Bench task taxonomy and evaluation dimensions. The six families are organized by task-defining physical state changes rather than semantic activity labels; M1–M5 cover task realization, source-relative interaction, target embodiment, and presentation quality.
ElementEmbodied-manipulation concernH2R-Bench operationalizationDesign origin
Task families: task-defining physical state changes
F1: Rigid rearrangementObject pose, support, or containmentPlace or transport a rigid object into the demonstrated spatial relation.Rigid transport and placement tasks.
F2: Mechanism actuationState of an articulated mechanismOpen, close, toggle, press, or rotate a task-relevant mechanism.Interaction with articulated objects.
F3: Insertion and assemblyConnection, fit, or attachment relationEstablish or remove a constrained connection between entities.Precision alignment and constrained contact.
F4: Deformable configurationNon-rigid shape or configurationProduce the demonstrated fold, bend, compression, or shape change.Deformable-object manipulation.
F5: Bulk-material transferDistribution or containment of materialPour, scoop, transfer, or mix material between regions or containers.Many-particle and material-flow manipulation.
F6: Surface/material transformationLocal surface condition or material integrityProduce a visible local change through wiping, spreading, cutting, peeling, or related interaction.Tool-mediated local transformation.
Evaluation dimensions: evidence required for valid H2R generation
M1: Goal-State CompletionWas the demonstrated task state reached?Score weighted predicates over the visible final state.Task-success and outcome evaluation in robotic video benchmarks.
M2: Action-Event CompletionWere the required manipulation events shown?Score completion of source-derived action events.Action-completeness evaluation in robotic video benchmarks.
M3: Functional Contact TransferDid robot contact support the source-consistent object response?Evaluate contact region, establishment, mode, temporal object response, and embodiment-compatible strategy.Contact-mediated manipulation and H2R interaction grounding.
M4: Embodiment CorrectnessDid the requested robot perform the manipulation?Evaluate robot presence, human absence, embodiment category, end-effector subtype, and temporal structure.Robot-structure and embodiment compliance.
M5: Video QualityIs the generated video visually and temporally well formed?Measure imaging quality, aesthetic quality, temporal stability, and motion smoothness.Task-agnostic video-generation quality.
Table S2: Dataset construction and annotation quality controls. Source-side annotations define the benchmark specification; detailed frame-level evidence and uncertainty fields are not exposed to generation models.
StageOutputQuality control
Source curationCurated clips across six familiesIdentifiable manipulated entities, visible task-relevant state change, and sufficient interaction evidence.
Clip grounding32 ordered source frames per clipFirst and last frames are retained; annotations identify initial state, contact, core transition, and final state when visible.
Structured annotationGoals, events, milestones, interaction, and embodiment strategiesStrict JSON schema with canonical event identifiers and a fixed six-family taxonomy.
Completion and validationSchema-complete case recordsOnly missing fields are completed; existing non-empty fields are preserved; schema and type mismatches are rejected.
Manual verificationAll benchmark annotationsEvery record is checked against its source video and corrected before evaluation.
Uncertainty accountingValidity and ambiguity fieldsEvery record has a validity decision; notes document residual uncertainty without changing the official evaluation set.
Table S3: Matched source-conditioning ablation for Seedance 2.0 on 24 sources. Both target embodiments are evaluated for every source. “Video” uses the full source clip; “9 ordered images” uses uniformly sampled chronological frames.
TargetInputM1M2M3M4M5Core
Parallel-Jaw GripperVideo0.7370.8500.7090.5980.79370.9
9 images0.7440.8480.2840.1040.78343.4
Dexterous HandVideo0.7680.8570.8610.8980.79985.1
9 images0.7530.8560.3130.0790.78243.7
Table S4: Main-evaluation H2RCore point estimates and percentile 95% confidence intervals. Intervals summarize variation across source tasks under the five-metric aggregation rule.
ModelParallel-Jaw Gripper Core [95% CI]Dexterous Hand Core [95% CI]
Kling-V374.5 [73.6, 75.3]81.7 [80.8, 82.6]
Seedance 2.077.3 [76.5, 78.1]84.6 [84.0, 85.3]
Wan2.776.5 [75.7, 77.4]83.1 [82.3, 83.9]
Grok Imagine50.1 [49.0, 51.1]49.2 [48.3, 50.1]
Veo 3.149.6 [48.9, 50.4]57.0 [56.1, 57.8]
Hunyuan-1.530.0 [28.8, 31.2]30.7 [29.5, 31.8]
LTX-2.332.1 [31.0, 33.2]39.8 [38.6, 40.9]
Mitty-EPIC61.5 [60.5, 62.5]56.1 [55.2, 57.0]
SkyReels-V331.5 [30.5, 32.6]34.6 [33.6, 35.6]
LongCat31.0 [29.9, 32.1]32.0 [30.9, 33.2]
Wan2.232.4 [31.4, 33.3]33.7 [32.7, 34.7]
Table S5: Paired source-task H2RCore differences for the leading video-conditioned models. Positive values favor the left model; percentile 95% CIs pair the same source tasks.
ComparisonGripper Δ Core [95% CI]Hand Δ Core [95% CI]
Seedance − Wan2.7+0.79 [+0.08, +1.48]+1.52 [+0.92, +2.16]
Seedance − Kling+2.82 [+2.04, +3.61]+2.89 [+2.15, +3.73]
Wan2.7 − Kling+2.04 [+1.40, +2.66]+1.37 [+0.66, +2.16]
Table S6: Human–automatic agreement for paired human and MLLM evaluations. M1–M4 are compared per video; the transfer subscore combines these four metrics without M5.
ScoreM1M2M3M4Transfer Subscore
Pearson r0.7910.8180.8800.8770.930
Table S7: Pairwise inter-judge agreement on the 120-source main evaluation set. Each cell reports Pearson correlation r between per-video scores.
Judge PairM1M2M3M4
Gemini / Qwen0.6950.6910.7410.874
Gemini / GPT0.5820.6360.7260.891
Qwen / GPT0.6560.6900.8090.892
Table S8: Source-conditioning interfaces in the main experiment. Frame inputs are sampled in temporal order from the source demonstration. No model receives a target-robot reference image.
ModelInterfaceSource input
Seedance 2.0VideoFull clip
Wan2.7VideoFull clip
Kling-V3VideoFull clip
Veo 3.1FrameFirst + last
Grok Imagine VideoFrameUp to 7 frames
HunyuanVideo 1.5-I2VFrameOrdered frame(s)
LTX-2.3FrameOrdered frame(s)
Mitty-EPIC14BFrameOrdered frame(s)
SkyReels-V3-R2VFrameOrdered frame(s)
LongCatFrameOrdered frame(s)
Wan2.2FrameFirst frame
Table S9: Native decoded output profiles in the main experiment. “Frames @ fps” reports decoded frame count and frame rate. Wan2.2 preserves source aspect ratio within its 720P-area setting; audio is removed before evaluation.
ModelDecoded SizeFrames @ fps
Seedance 2.01280×720121 @ 24
Wan2.71280×720150 @ 30
Kling-V31280×720121 @ 24
Veo 3.11280×720144 @ 24
Grok Imagine Video1280×720121 @ 24
HunyuanVideo 1.5-I2V1280×720121 @ 24
LTX-2.31280×720120 @ 24
Mitty-EPIC14B1280×72037 @ 8
SkyReels-V3-R2V1280×720121 @ 24
LongCat832×48080 @ 15
Wan2.2720P area121 @ 24
Table S10: Main-evaluation metric breakdown by task family for all 11 evaluated models. Parallel-Jaw Gripper and Dexterous Hand targets are reported separately. Video Quality is the task-family mean of M5 and contributes 0.10 to H2RCore. Overall values in Table 2 are computed from the unrounded per-video scores.
Parallel-Jaw GripperDexterous Hand
TaskModelGoal CompletionAction CompletionContact TransferEmbod. Correct.Video QualityCoreGoal CompletionAction CompletionContact TransferEmbod. Correct.Video QualityCore
Kling-V30.7360.8370.7820.7350.80377.10.7380.8370.8470.8860.80783.7
Seedance 2.00.7620.8570.8360.8240.79882.10.7750.8530.8800.9080.80286.1
Wan2.70.7440.8430.8240.7760.79379.70.7770.8450.8900.9130.79986.4
Grok Imagine0.6080.7040.5550.4300.80657.30.6560.7420.5610.3680.81157.0
Veo 3.10.7510.7910.5740.1970.79054.20.7800.8300.6530.3050.80060.9
Hunyuan-1.50.5200.4460.1880.0000.81228.20.4830.4290.2090.0270.81428.9
LTX-2.30.4240.5500.3440.0180.78433.30.4550.5580.4220.1700.79140.9
Mitty-EPIC0.5220.5910.5400.5750.73357.50.5280.6400.5910.4020.73354.6
SkyReels-V30.4080.5270.2670.0000.79229.90.3760.4890.3330.0610.79532.8
LongCat0.4120.5020.2080.0000.79927.90.3940.4760.3060.0320.79831.2
F1 RigidWan2.20.3770.5470.2290.0000.77628.50.4220.5530.2490.0000.77429.8
Kling-V30.6410.7920.6860.6820.80270.60.6660.7960.7700.8850.80479.6
Seedance 2.00.7600.8470.7370.7420.79776.40.7350.8540.8280.9270.80184.5
Wan2.70.6950.7830.7040.7710.79474.40.6760.7950.7780.9130.79980.8
Grok Imagine0.6460.7290.4910.5270.79159.10.6830.7700.5730.3440.80757.4
Veo 3.10.6700.7540.5080.0490.78545.90.6910.8240.6200.1510.79353.8
Hunyuan-1.50.4330.5170.1780.0140.80728.10.3990.4810.2020.1230.80931.0
LTX-2.30.4780.5140.2780.0000.77431.00.5230.6130.3280.1570.78639.4
Mitty-EPIC0.5150.5880.5370.5790.74657.50.5070.6090.5400.3930.74652.2
SkyReels-V30.4540.5520.2230.0000.79029.70.4250.5590.3030.0110.79432.1
LongCat0.4540.5790.2430.0000.79230.70.4030.5770.2750.0110.79431.2
F2 MechanismWan2.20.4270.6070.2340.0000.76330.20.5260.6720.2640.0000.76233.5
Kling-V30.6830.7480.7480.7350.79473.90.6800.7390.7960.9070.79980.4
Seedance 2.00.7220.7920.8340.8390.78980.80.7370.8320.8670.9070.79384.7
Wan2.70.6230.7270.7420.7850.78973.90.6600.7620.8130.9120.79281.0
Grok Imagine0.5990.6880.4220.3200.78849.40.6030.6500.3980.2990.80747.8
Veo 3.10.6870.7640.5040.0720.78046.90.6190.7700.5970.2400.79153.8
Hunyuan-1.50.4620.4480.2140.0000.80328.10.4090.4690.1930.0220.80527.7
LTX-2.30.4260.4700.2760.0460.77330.80.5100.5160.3430.1950.77939.3
Table S11: Metric implementation summary. Detailed definitions, evidence rules, diagnostic components, and per-video formulas are given in Section B.6.
MetricEvidenceComputation
M1: Goal25 generated frames and final-state predicates.Weighted mean of normalized 0–4 predicate scores.
M2: Action25 generated frames and required events.Weighted mean of normalized event-completion scores.
M3: Contact25 source frames, 25 generated frames, and the contact specification.Mean of applicable contact dimensions; zero when source grounding fails.
M4: Embodiment25 generated frames and the target specification.Weighted mean of five embodiment dimensions; zero on a hard failure.
M5: QualityDecoded generated video.Mean of MUSIQ, CLIP aesthetic, temporal-stability, and AMT-S scores.

실제로 확인된 결과

  • 원본 영상 전체를 입력받는 3개 모델(Seedance 2.0, Wan2.7, Kling-V3)이 상위권을 차지했고, Seedance 2.0이 H2RCore 그리퍼 77.3점, 다섯손가락 손 84.6점으로 1위, Wan2.7이 76.5/83.1점, Kling-V3가 74.5/81.7점으로 뒤를 이었다.
  • 영상 품질(M5)은 모델 간 0.73~0.81로 거의 차이가 없었지만 H2RCore는 30.0~84.6점까지 크게 벌어졌으며 둘의 순위 상관은 약했다(스피어만 ρ=0.14).
  • 사람 평가자 3명이 채점한 결과와 AI 판정관 점수의 순위 상관은 ρ=0.883, 종합 상관은 피어슨 r=0.930으로 높게 나와 자동 평가가 사람 평가와 상당히 일치했다.
  • 11개 모델 중 9개가 평행 그리퍼보다 다섯손가락 로봇 손 조건에서 평균 3.3점 더 높은 점수를 받았고, 영상 조건 입력 모델 3종(Seedance, Wan2.7, Kling-V3)에서는 각각 7.3, 6.6, 7.2점 상승으로 효과가 더 크게 나타났다.
  • 로봇 참고 이미지를 추가로 준 경우 Wan2.7은 H2RCore가 76.5에서 83.1로 크게 올랐지만 Kling-V3는 13.5점, Seedance는 8.8점 하락해 모델마다 효과가 정반대로 나타났다.

어디에 쓸 수 있나

  • 사람이 찍은 작업 영상을 로봇 학습 데이터로 자동 변환하려는 파이프라인을 만들 때, 완성된 후 영상 품질만 보지 말고 접촉·형태 일치 여부를 별도로 점검하는 체크리스트로 활용할 수 있다.
  • 영상 생성 AI를 로봇 데이터 증강 도구로 도입하기 전에, 이 벤치마크의 5가지 평가 축(목표·동작·접촉·형태·품질)을 자체 평가 기준으로 참고할 수 있다.
  • 그리퍼용과 다섯손가락 손용 중 어떤 로봇 형태를 지정할지 고를 때, 사람 손 모양과 더 비슷한 형태를 쓰는 쪽이 접촉 전이에 유리할 수 있다는 참고 자료로 쓸 수 있다.

한계와 남은 검증

  • 이 벤치마크는 영상에 보이는 증거만 평가하며, 실제 물리적으로 로봇이 그 동작을 수행할 수 있는지나 학습된 정책의 실제 성능은 확인하지 않는다.
  • EgoDex에서 뽑은 120개 짧은 클립과 그리퍼·다섯손가락 손 두 형태만 다루므로 로봇 종류나 작업 상황의 다양성은 일부만 반영한다.
  • 각 모델이 지원하는 입력 방식(전체 영상 vs 몇 장의 프레임)이 서로 달라 모델 능력과 인터페이스 차이가 섞여 있어 순수한 성능 비교로 보기는 어렵다.
  • 가려짐, 미세한 접촉, 심한 생성 오류가 있는 경우 AI 판정관의 채점도 불확실할 수 있으며, 더 긴 시범 영상이나 3차원 정보를 활용한 확장은 향후 과제로 남아 있다.

왜 중요한가

로봇 학습용 데이터를 만들 때 '사람 영상을 로봇 영상으로 바꿔주는 AI'를 쓰겠다는 아이디어가 유행하는데, 이 연구는 지금 나온 모델들이 겉보기 좋은 영상만 만들고 실제로는 로봇이 물건을 제대로 만지지도, 요청한 로봇 형태를 유지하지도 못한다는 걸 구체적 수치로 보여준다. 앞으로 이런 영상 생성 모델을 로봇 데이터로 쓰려는 사람들은 화질이나 그럴듯함이 아니라 접촉과 형태 일치를 반드시 따로 확인해야 한다.

이 논문의 용어

  • world model(월드 모델) · 미래 상황이나 영상을 예측해 만들어내는 AI 모델. 여기서는 영상 생성 AI를 로봇의 미래 행동을 예측하는 도구로 본다
  • 임바디먼트(embodiment) · 로봇의 몸 형태, 예를 들어 평행 그리퍼나 다섯손가락 로봇 손처럼 실제 물체를 다루는 신체 구조
  • H2RCore · 다섯 가지 평가 항목(M1~M5)을 가중평균해 0~100점으로 합친 이 연구의 종합 점수
  • MLLM 판정관 · 영상을 보고 점수를 매기는 대형 멀티모달 언어모델. 이 논문에서는 Gemini, Qwen, GPT 세 개를 사용했다
  • functional contact(기능적 접촉) · 로봇이 물체와 실제로 작업 목적에 맞게 접촉했는지를 뜻하며, 손 모양이나 경로가 달라도 같은 기능을 하면 인정한다

저자 · Dingyi Rong

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL MEDIA 최신 기사

그림 출처: Dingyi Rong et al., arXiv:2608.13049, arxiv-nonexclusive