AI가 만든 동영상이 '결과'를 실제로 이뤄냈는지 채점하는 벤치마크가 나왔다
AI가 만든 동영상이 '결과'를 실제로 이뤄냈는지 채점하는 벤치마크가 나왔다
이 논문은 참조 이미지를 보고 지시받은 작업을 '결과 중심'으로 완수하는지 평가하는 새로운 과제와 데이터셋 SemComp-Data, 평가 도구 SemComp-Bench를 만들었다. 예를 들어 지폐 사진을 주고 '거북이 모양으로 접어라'라고 하면, 접는 과정은 안 보여줘도 되지만 그 지폐가 실제로 거북이 모양이 되어야 성공으로 친다. 실제 비디오 생성 모델들을 테스트해본 결과, 가장 잘하는 모델도 결과 달성 점수가 40%를 넘지 못했다.
METAL MEDIA 해설 도표
AI가 만든 동영상이 '결과'를 실제로 이뤄냈는지 채점하는 벤치마크가 나왔다
- 01동영상 생성 AI가 중간 과정을 그럴듯하게 보여주는 데는 강하지만, 정작 지시받은 '결과'를 이미지와 의미적으로 연결지어 완성하는지는 제대로 평가된 적이 없었다는 문제의식에서 출발했다.
- 02실제 유튜브류 영상(Koala-36M 데이터셋)에서 후보 영상 거르기, 시작 상태와 결과 상태 찾아내기, 결과 중심 클립 추출하기, 간단/상세 두 버전의 지시문 만들기라는 4단계 파이프라인을 거쳐 1,273개의 '이미지-지시문-영상' 세트를 자동으로 만들었다.
- 03이 중 60개를 고르게 뽑은 SemComp-Core로 Wan2.2, HunyuanVideo, CogVideoX 등 대표 영상 생성 모델 7종을 테스트했으며, 채점은 사람 대신 영상언어모델(VLM)이 27개 프레임을 보고 예/아니오로 답하는 방식을 썼다.
- 04평가는 크게 두 축으로 나뉜다. '결과 달성(OA)'은 지시한 결과가 실제로 나왔는지, 참조 이미지와 의미적으로 맞는지를 보고, '생성 신뢰성(GR)'은 물리적으로 말이 되는지, 화면이 깨지지 않았는지를 본다.
- 05가장 성적이 좋았던 HunyuanVideo-1.5도 결과 달성 점수가 37.8%에 그쳤고, 참조 이미지 없이 텍스트만으로 만드는 방식(T2V)보다 이미지를 함께 준 방식(I2V)이 훨씬 나은 것으로 나타나 참조 이미지 조건화의 중요성이 확인됐다.
무엇을 했나
- 동영상 생성 AI가 중간 과정을 그럴듯하게 보여주는 데는 강하지만, 정작 지시받은 '결과'를 이미지와 의미적으로 연결지어 완성하는지는 제대로 평가된 적이 없었다는 문제의식에서 출발했다.
- 실제 유튜브류 영상(Koala-36M 데이터셋)에서 후보 영상 거르기, 시작 상태와 결과 상태 찾아내기, 결과 중심 클립 추출하기, 간단/상세 두 버전의 지시문 만들기라는 4단계 파이프라인을 거쳐 1,273개의 '이미지-지시문-영상' 세트를 자동으로 만들었다.
- 이 중 60개를 고르게 뽑은 SemComp-Core로 Wan2.2, HunyuanVideo, CogVideoX 등 대표 영상 생성 모델 7종을 테스트했으며, 채점은 사람 대신 영상언어모델(VLM)이 27개 프레임을 보고 예/아니오로 답하는 방식을 썼다.
- 평가는 크게 두 축으로 나뉜다. '결과 달성(OA)'은 지시한 결과가 실제로 나왔는지, 참조 이미지와 의미적으로 맞는지를 보고, '생성 신뢰성(GR)'은 물리적으로 말이 되는지, 화면이 깨지지 않았는지를 본다.
- 가장 성적이 좋았던 HunyuanVideo-1.5도 결과 달성 점수가 37.8%에 그쳤고, 참조 이미지 없이 텍스트만으로 만드는 방식(T2V)보다 이미지를 함께 준 방식(I2V)이 훨씬 나은 것으로 나타나 참조 이미지 조건화의 중요성이 확인됐다.
| Models | Aor | Asg | Agec | Agvc | OA Score |
|---|---|---|---|---|---|
| Seedance 2.0 [5] | 0.839 | 0.744 | 0.444 | 0.594 | 20.0% |
| Wan2.2-TI2V-5B [28] | 0.589 | 0.400 | 0.689 | 0.922 | 23.3% |
| Wan2.2-I2V-A14B [28] | 0.800 | 0.528 | 0.628 | 0.789 | 28.3% |
| CogVideoX1.5-5B-I2V [30] | 0.550 | 0.389 | 0.506 | 0.744 | 14.4% |
| SkyReels-V2-I2V-14B-720P [7] | 0.733 | 0.489 | 0.522 | 0.772 | 22.8% |
| HY†-1.5-720P-I2V | 0.878 | 0.706 | 0.583 | 0.794 | 37.8% |
| Phantom-1.3B [20] | 0.539 | 0.356 | 0.322 | 0.511 | 3.9% |
| Models | Gpp | Gvc | Gafr | Gwsc | Gti | GR Score |
|---|---|---|---|---|---|---|
| Seedance 2.0 [5] | 0.883 | 0.994 | 0.994 | 0.739 | 0.978 | 91.8% |
| Wan2.2-TI2V-5B [28] | 0.794 | 0.978 | 0.889 | 0.722 | 0.889 | 85.4% |
| Wan2.2-I2V-A14B [28] | 0.911 | 0.972 | 0.967 | 0.672 | 0.928 | 89.0% |
| CogVideoX1.5-5B-I2V [30] | 0.944 | 0.811 | 0.772 | 0.583 | 0.828 | 78.8% |
| SkyReels-V2-I2V-14B-720P [7] | 0.778 | 0.922 | 0.739 | 0.328 | 0.789 | 71.1% |
| HY†-1.5-720P-I2V | 0.800 | 0.939 | 0.883 | 0.472 | 0.861 | 79.1% |
| Phantom-1.3B [20] | 0.778 | 0.972 | 0.856 | 0.506 | 0.728 | 76.8% |
| Model | Modality | Instruction Type | Aor | Asg | Agec | Agvc | OA Score |
|---|---|---|---|---|---|---|---|
| Wan2.2-A14B [28] | I2V | Detailed | 0.800 | 0.528 | 0.628 | 0.789 | 28.3% |
| T2V | Detailed | 0.889 | 0.522 | 0.317 | 0.117 | 4.4% | |
| T2V | Brief | 0.389 | 0.111 | 0.806 | 0.250 | 0.6% | |
| CogVideoX1.5-5B [30] | I2V | Detailed | 0.550 | 0.389 | 0.506 | 0.744 | 14.4% |
| T2V | Detailed | 0.567 | 0.272 | 0.361 | 0.161 | 5.0% | |
| T2V | Brief | 0.567 | 0.150 | 0.678 | 0.161 | 0.6% | |
| HY†-1.5-720P | I2V | Detailed | 0.878 | 0.706 | 0.583 | 0.794 | 37.8% |
| T2V | Detailed | 0.833 | 0.606 | 0.389 | 0.133 | 4.4% | |
| T2V | Brief | 0.550 | 0.100 | 0.761 | 0.178 | 1.7% |
| Configuration Group | Keyword Strings |
|---|---|
| news_keywords (13) | News; Briefing; Coverage; Current Affairs; Deep Dive; Documentary; Exclusive; Headline; Interview; Live; Live Update; Press Conference; Report. |
| movie_keywords (13) | Behind the scenes; Blooper; Casting; Movie Clip; Playback; Preview; Review; Spoiler; Teaser; TV Series; Talk; Talks; TalkShow. |
| entertainment_keywords (19) | ASMR; Celebrity; Challenge; Concert; Gaming; Highlights; Music Video; Podcast; Prank; Reaction; Reacts; Travel; Unboxing; Variety Show; Vlog; HBO; SportsCenter; NFL; NBA. |
| Domain | Category | Definition |
|---|---|---|
| Food and Cooking | Dish Making | A dish or food product is completed from ingredients. |
| Food Transformation | The same food item undergoes a visible cooking-related state change. | |
| Food Plating | Food elements are arranged into a complete presentation. | |
| Beauty and Fashion | Styling | The appearance of the same person is visibly restyled. |
| Tool Cleaning | A beauty or fashion tool is cleaned to a reusable condition. | |
| Try-on | An external appearance reference is applied to a subject. | |
| Product Reveal | A previously concealed product becomes visible. | |
| Sports and Fitness | Gear Setup | Sports gear or equipment is configured into a ready-to-use state. |
| Body Preparation | The subject’s body enters a sports-related prepared state. | |
| Sports Edits | Source event footage or identity cues are transformed into a completed sports-media edit. | |
| Crafts and DIY | Assembly | Separate parts are combined into a complete structure or object. |
| Restoration | The condition of the same object is restored or refurbished. | |
| Renovation | A built space is reorganized or remodeled into a new spatial state. | |
| Blueprint Construction | A physical structure is built according to an abstract design specification. | |
| Material Shaping | A raw material is physically reshaped into a new form. | |
| Gardening and Pets | Plant Treatment | A plant is improved through care-related operations. |
| Floral Design | Floral or plant elements are trimmed and arranged into an aesthetic design. | |
| Pet Grooming | A pet’s appearance is improved through grooming. | |
| Arts and Precision | Artwork Creation | An artwork is completed from an unfinished visual basis. |
| Digital Creation | A digital creative product is completed from a visual or abstract reference. | |
| Sculpting | A sculptable material is turned into a finished three-dimensional artwork. |
| Category | Reference State | Outcome State |
|---|---|---|
| Dish Making | Ingredients or an incomplete food preparation are visible. | A recognizable completed dish or food product is visible. |
| Food Transformation | A food item is visible before a cooking-related state change. | The same food item exhibits the intended visible state change. |
| Food Plating | Food components are present before final arrangement. | The food components form a complete plated presentation. |
| Styling | A person is visible before the target styling change. | The same person visibly exhibits the completed styling result. |
| Tool Cleaning | A beauty or fashion tool is visibly soiled or not ready for reuse. | The same tool is visibly clean and reusable. |
| Try-on | The subject and an external appearance reference are identifiable. | The subject visibly exhibits the referenced garment, accessory, or appearance. |
| Product Reveal | The product is concealed, covered, or not yet identifiable. | The product is exposed and visually identifiable. |
| Gear Setup | Sports gear is unconfigured, disassembled, or not ready for use. | The gear is visibly configured in a ready-to-use state. |
| Body Preparation | A subject is visible before completing a sports-related preparation. | The subject visibly reaches the intended prepared state. |
| Sports Edits | Source event footage or identity cues for a sports edit are visible. | A completed sports-media edit is visible. |
| Assembly | Separate or partially combined components are visible. | The components form a complete structure or object. |
| Restoration | A worn, damaged, or degraded object is visible. | The same object is visibly restored or refurbished. |
| Renovation | A built space is visible before reorganization or remodeling. | The space exhibits the completed renovated layout or appearance. |
| Blueprint Construction | An abstract design, plan, or blueprint is visible. | A corresponding completed physical structure is visible. |
| Material Shaping | Raw or incompletely shaped material is visible. | The material has the intended completed form. |
| Plant Treatment | A plant is visible before care or treatment. | The plant exhibits a visibly improved state after treatment. |
| Floral Design | Unarranged floral or plant elements are visible. | The elements form a completed aesthetic arrangement. |
| Pet Grooming | A pet is visible before grooming. | The same pet exhibits the completed grooming result. |
| Artwork Creation | An unfinished artwork or visual basis is visible. | A completed artwork is visible. |
| Digital Creation | A visual reference or unfinished digital artifact is visible. | A completed digital creative product is visible. |
| Sculpting | Raw or partially shaped sculptable material is visible. | A completed three-dimensional sculpture is visible. |
| Alignment Type | Operational Definition | Typical Relevant Attributes |
|---|---|---|
| Object Element | Grounds the outcome in task-relevant object categories, components, ingredients, materials, quantities, or colors appearing in the reference, without requiring the exact appearance of a specific object instance. | Object category, composition, count, material, and color palette. |
| Identity | Requires the depicted person to remain identifiable across the reference and generated outcome. | Person identity, face appearance, body appearance, and pose. |
| Object Appearance | Grounds the outcome in the instance-level visual appearance of a specific reference object. | Object category, color palette, material, shape structure, surface appearance, and graphic details. |
| Scene | Grounds the outcome in the organization or context of the reference scene. | Scene type, scene layout, spatial relation, background context, and viewpoint. |
| Complete preservation-attribute vocabulary (17): object_category, object_composition, object_count, material, color_palette, shape_structure, surface_appearance, graphic_details, person_identity, face_appearance, body_appearance, pose, scene_type, scene_layout, spatial_relation, background_context, and viewpoint. |
| Model | σ(Aor) | σ(Asg) | σ(Agec) | σ(Agvc) | σ(OA Score) |
|---|---|---|---|---|---|
| Seedance 2.0 | 0.96 | 0.96 | 2.55 | 7.52 | 4.41 |
| Wan2.2-TI2V-5B | 4.19 | 4.41 | 3.85 | 0.96 | 3.33 |
| Wan2.2-I2V-A14B | 1.67 | 2.55 | 5.09 | 3.47 | 2.89 |
| CogVideoX1.5-5B-I2V | 3.33 | 4.19 | 2.55 | 0.96 | 0.96 |
| SkyReels-V2-I2V-14B-720P | 1.67 | 3.47 | 3.47 | 2.55 | 0.96 |
| HY†-1.5-720P-I2V | 2.55 | 2.55 | 2.89 | 1.92 | 3.47 |
| Phantom-1.3B | 0.96 | 1.92 | 1.92 | 0.96 | 0.96 |
| Model | σ(Gpp) | σ(Gvc) | σ(Gafr) | σ(Gwsc) | σ(Gti) | σ(GR Score) |
|---|---|---|---|---|---|---|
| Seedance 2.0 | 3.33 | 0.96 | 0.96 | 1.92 | 0.96 | 1.58 |
| Wan2.2-TI2V-5B | 0.96 | 2.55 | 6.74 | 12.95 | 0.96 | 3.95 |
| Wan2.2-I2V-A14B | 5.85 | 0.96 | 1.67 | 1.92 | 4.19 | 1.53 |
| CogVideoX1.5-5B-I2V | 2.55 | 13.47 | 6.31 | 8.82 | 8.55 | 7.12 |
| SkyReels-V2-I2V-14B-720P | 0.96 | 3.47 | 6.74 | 11.10 | 3.85 | 4.53 |
| HY†-1.5-720P-I2V | 8.82 | 1.92 | 0.00 | 6.94 | 0.96 | 1.35 |
| Phantom-1.3B | 5.09 | 0.96 | 8.22 | 15.84 | 6.94 | 7.20 |
| Model | Modality | Instruction | σ(Aor) | σ(Asg) | σ(Agec) | σ(Agvc) | σ(OA Score) |
|---|---|---|---|---|---|---|---|
| Wan2.2-A14B | I2V | Detailed | 1.67 | 2.55 | 5.09 | 3.47 | 2.89 |
| T2V | Detailed | 2.55 | 4.19 | 1.67 | 4.41 | 2.55 | |
| T2V | Brief | 0.96 | 0.96 | 2.55 | 7.26 | 0.96 | |
| CogVideoX1.5-5B | I2V | Detailed | 3.33 | 4.19 | 2.55 | 0.96 | 0.96 |
| T2V | Detailed | 3.33 | 0.96 | 2.55 | 0.96 | 1.67 | |
| T2V | Brief | 2.89 | 1.67 | 0.96 | 3.47 | 0.96 | |
| HY†-1.5-720P | I2V | Detailed | 2.55 | 2.55 | 2.89 | 1.92 | 3.47 |
| T2V | Detailed | 0.00 | 2.55 | 6.74 | 5.00 | 1.92 | |
| T2V | Brief | 1.67 | 3.33 | 4.19 | 5.36 | 0.00 |
왜 중요한가
지금까지의 동영상 생성 평가는 화질이나 움직임의 자연스러움을 주로 봤지만, 실제 활용에서는 '시킨 일을 진짜로 해냈는가'가 더 중요한 질문이다. 이 벤치마크는 그 질문에 답할 수 있는 기준을 처음으로 체계화했다는 점에서, 앞으로 작업 지시형 영상 생성 모델을 만들거나 고를 때 참고할 수 있는 잣대가 된다.
이 논문의 용어
- VLM(Vision-Language Model) · 이미지나 영상을 보고 텍스트로 답할 수 있는 AI 모델
- I2V / T2V · 이미지+텍스트로 영상을 만드는 방식(I2V) vs 텍스트만으로 만드는 방식(T2V)
- OA Score / GR Score · 결과를 제대로 달성했는지 점수(OA), 영상이 물리적·시각적으로 깨지지 않고 신뢰할 만한지 점수(GR)
- 시맨틱 그라운딩(Semantic Grounding) · 생성된 결과가 원본 참조 이미지와 의미상 제대로 연결되어 있는지의 정도
- Koala-36M · 실제 세계를 촬영한 대규모 동영상 원본 데이터셋
본문에 싣지 못한 그림
- Figure 1: Demonstration of Semantic Task Completion Video Generation, which focuses not on the transformation process but solely on outcome achievement and semantic grounding.
- Figure 2: Representative SemComp-Data instances. Each example shows a reference frame, its outcome-centric video clip, and the brief instruction for compact presentation; every instance also includes a detailed instruction.
- Figure 3: Overview of the four-stage SemComp-Data curation pipeline. Panels (a)–(d) show Candidate Filtering, State Mining, Video Extension, and Instruction Structuring, respectively. The right column shows the corresponding stage outputs, and the bottom row presents the final data triplet.
- Figure 4: Statistics of SemComp-Data: (a) distribution of alignment types across domains, with distinct textures denoting different types; (b) high-frequency words in the instructions; and (c) distribution of evaluation instances across the six domains.
- Figure 5: Generated-video examples. The outcome-centric video clip and the reference image are extracted from the same full-context video.
- Figure S1: Common system prompt used for SemComp-Bench evaluation.
- Figure S2: Outcome Achievement prompt for Outcome Realization (QOA1) and Semantic Grounding (QOA2).
- Figure S3: Outcome Achievement prompt for Grounded Entity Consistency (QOA3) and Global Visual Continuity (QOA4).
- Figure S4: Generation Reliability prompt for Physical Plausibility (QGR1).
- Figure S5: Generation Reliability prompt for Visual Clarity, Artifact-Free Rendering, Within-Scene Spatiotemporal Coherence, and Text and Interface Integrity (QGR2–QGR5).
- Figure S7: Additional generated-video comparisons on four SemComp-Core instances. Outputs from all seven I2V models in each block were generated using the shared reference condition and detailed instruction; the paired brief instruction is displayed only for compact presentation. Temporally ordered output frames illustrate differences in both outcome achievement and generation reliability.
최신 논문
- AI 코딩 에이전트에게 과학 소프트웨어 수리를 시켜보니, 절반도 제대로 못 고쳤다AI 코딩 에이전트에게 과학 소프트웨어 수리를 시켜보니, 절반도 제대로 못 고쳤다
- 논문 속 시연이 아니라 실제 서비스에 넣을 수 있는 희소 어텐션 만들기논문 속 시연이 아니라 실제 서비스에 넣을 수 있는 희소 어텐션 만들기
- 고객상담 AI 상담원이 규정을 '한 번의 행동'이 아니라 '전체 절차'로 지키게 만드는 방법고객상담 AI 상담원이 규정을 '한 번의 행동'이 아니라 '전체 절차'로 지키게 만드는 방법
- 로봇 팔에게 사람의 시연 없이 새 일 시키기, 말 잘하는 AI가 대신 가르친다로봇 팔에게 사람의 시연 없이 새 일 시키기, 말 잘하는 AI가 대신 가르친다
- 에이전트 학습용 환경을 새로 만드는 대신, 기존 환경에 '패치 부품'을 씌워 그 에이전트의 약점에 맞게 바꾸는 방법에이전트 학습용 환경을 새로 만드는 대신, 기존 환경에 '패치 부품'을 씌워 그 에이전트의 약점에 맞게 바꾸는 방법
- AI 모델을 '소유'하지 못한 조직은 안전 통제도 절반밖에 못 한다AI 모델을 '소유'하지 못한 조직은 안전 통제도 절반밖에 못 한다
- AI가 선생님 모델을 따라 배우다가, 정답에 다가가는 '좋은 생각'까지 억누르는 문제를 잡아낸다AI가 선생님 모델을 따라 배우다가, 정답에 다가가는 '좋은 생각'까지 억누르는 문제를 잡아낸다
- AI가 특정 사람 말투를 흉내내도록 시켜봤더니, 결국 AI 자신의 말투에서 못 벗어난다AI가 특정 사람 말투를 흉내내도록 시켜봤더니, 결국 AI 자신의 말투에서 못 벗어난다