K-文化的一切——从回归到 K-美妆,发送到您的邮箱订阅邮件

METAL MEDIA

AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities

arXiv:2607.248212026-07-16

终于有一个基准能检验视频编辑AI是不是把声音也一起改对了,作者还顺手给出了做得更好的智能体

基于指令的视频编辑模型常常把画面改得很像样,却把声音改错甚至误删。AVE-Compass是一个诊断型基准,包含145段原始视频、196条音画耦合的编辑指令和2688条检查项,专门用来抓这类跨模态失误。作者据此发现的问题,进一步设计出按“规划-执行-评估”三步走的AVE-Agent,并与现有六个编辑系统做了对比。

METAL MEDIA 解读图

AVE-Compass与AVE-Agent的协同运作方式

证据状态已报告实测结果

  1. 原始视频加编辑指令输入带原始音频的视频以及一句自然语言编辑请求
  2. 规划智能体分析视频内容,把指令拆解成视频/音频/语音子任务的依赖图,并为每步设定评估标准
  3. 执行智能体用对应工具执行各子任务,子任务评估器发现问题时自我检查并重试
  4. 综合评估智能体对合成后的完整片段重新评估指令遵循、保真度和质量,触发混音、子任务重做或整体重新规划
  5. AVE-Compass评分基于检查清单的指令遵循/保真分数、独立的真实感评分,以及唇形同步等自动指标共同诊断各模型的失败模式
这是 METAL MEDIA 制作的解读图,并非论文作者提供的原图。

他们做了什么

  1. 研究起点是现有视频编辑基准大多只评估无声视频的画面变化,或把音频编辑单独评测,无法覆盖真实场景中需要画面与声音协同修改的编辑需求。
  2. 团队整理了145段原始视频和196条经人工核验的音画耦合编辑指令,并让人工审核了2688条是/否检查项,分别核实编辑是否执行到位、未编辑内容是否被保留。
  3. 评测拆分为四个由多模态大模型评分的维度——指令遵循(IF)、保真保留(FP)、真实感(Realism),以及用IF乘以FP衡量“完整编辑成功”的编辑意图(EI)——再配合唇形同步、音画同步等七种自动化指标。
  4. 基于这些诊断结果,作者构建了AVE-Agent:规划智能体把指令拆成视频/音频/语音子任务的依赖图,执行智能体执行并自我检查纠错各子任务,评估智能体对合成后的完整片段重新打分并触发混音、重做或重新规划;并与Wan2.7、HappyHorse、Gemini-Omni、Seedance、LTX2等五个现有系统做了对比。
Figure 1: Illustration of AVE-Compass. Given the same editing instruction, the four panels illustrate characteristic ways current audio-visual editing models drift, with IF (Instruction Following) and FP (Fidelity Preserving) marks indicating where each output breaks.
Figure 1: Illustration of AVE-Compass. Given the same editing instruction, the four panels illustrate characteristic ways current audio-visual editing models drift, with IF (Instruction Following) and FP (Fidelity Preserving) marks indicating where each output breaks.
Table 1: Comparison of AVE-Compass with existing editing benchmarks. Target Modality denotes the evaluated streams. Task Labels summarize task coverage with normalized labels, including object, scene, motion, camera, VFX, audio event, and speech. # Metrics and # Task Categories count evaluation dimensions and editing operation types, respectively. Cross-Modal Evaluation indicates whether the benchmark evaluates audio–video coupled editing, ranging from object-level correspondence to global-scene dependencies with diagnostic checks. Multi-Shot and Speech Edit indicate support for multi-shot source clips and speech-related editing. Difficulty Analysis denotes stratified evaluation over factors such as target localization, audio-source complexity, and cross-modal linkage. AVE-Compass is the only benchmark that jointly supports all these capabilities.
BenchmarkTarget ModalityTargetModalityTask LabelsTaskLabels# Metrics#Metrics# Task Categories# TaskCategoriesCross-Modal EvaluationCross-ModalEvaluationMulti- ShotMulti-ShotSpeech Edit
Target
Modality
Task
Labels
#
Metrics
# Task
Categories
Cross-Modal
Evaluation
Multi-
Shot
Speech
Edit
Difficulty
Analysis
Video-Only Editing Benchmarks
VEBench Sun 2025VideoObject, Scene3N/A
IVEBench Chen 2025VideoObject, Scene1235
FiVE Li 2025VideoObject156
UniVBench Wei 2026VideoScene216
CoVEBench Wu 2026VideoObject, Scene1119
Audio-Visual Editing Benchmarks
SAVEBench Xu 2025VideoAudioObject121Object-level
AVED-Bench Lin 2026VideoAudioObject53Object-level
AVI-Edit Zheng 2025VideoAudioAudio Event73Object-level
AVE-Compass (Ours)VideoAudioObject, Scene, Audio Event, SpeechObject, Scene,Audio Event, Speech1928Global Scene
Object, Scene,
Audio Event, Speech
Figure 2: Dataset statistics of the AVE-Compass benchmark.
Figure 2: Dataset statistics of the AVE-Compass benchmark.
Table 2: MLLM-as-Judge results on AVE-Compass. We report four dimensions—Editing Intent, Instruction Following, Fidelity Preserving, and Realism—each as Overall / Video / Audio scores on a 0–100 scale. Models are ranked by the Overall Editing Intent score, which serves as the primary measure of complete edit execution by jointly accounting for instruction following and fidelity preservation. All metrics are higher-is-better. Best in bold. ∗Gemini misses 16 speech edits due to content moderation.
ModelEditing IntentInstruction FollowingFidelity PreservingRealism
OverallVideoAudioOverallVideoAudioOverallVideoAudioOverallVideoAudio
AVE-Agent (Wan)59.8+17.466.7+6.650.2+25.477.3+8.084.8+6.569.4+9.177.6+13.480.1+1.674.8+26.762.1+1.745.3+1.878.9+1.5
Wan2.742.460.124.869.378.360.364.278.548.160.443.577.4
HappyHorse41.356.718.866.975.554.563.974.553.363.049.976.0
Gemini-Omni∗38.056.110.044.974.010.884.976.896.166.949.284.5
Seedance26.636.113.537.450.524.481.783.080.469.254.084.4
LTX215.210.726.470.172.866.230.624.542.364.142.086.2
Figure 3: Overview of AVE-Agent. Given a source video with audio and an edit instruction, the planner agent first preprocesses the input audio-visual clip and analyzes it through structured captioning. Conditioned on this analysis, it decomposes the instruction into a dependency-aware subtask DAG with per-step evaluation criteria. The executor agent executes the subtasks in order. For each subtask, it refines the intent into tool prompts and tool specifications, routes the request to the video, audio, or speech branch, and optimizes the result through a reflection loop guided by the planner-provided criteria. Finally, the mixed evaluator agent consumes the temporary assembled clip and planner criteria, scores instruction following, Fidelity Preserving, and quality, and outputs a control signal for final pass, remixing, subtask regeneration, or replanning.
Figure 3: Overview of AVE-Agent. Given a source video with audio and an edit instruction, the planner agent first preprocesses the input audio-visual clip and analyzes it through structured captioning. Conditioned on this analysis, it decomposes the instruction into a dependency-aware subtask DAG with per-step evaluation criteria. The executor agent executes the subtasks in order. For each subtask, it refines the intent into tool prompts and tool specifications, routes the request to the video, audio, or speech branch, and optimizes the result through a reflection loop guided by the planner-provided criteria. Finally, the mixed evaluator agent consumes the temporary assembled clip and planner criteria, scores instruction following, Fidelity Preserving, and quality, and outputs a control signal for final pass, remixing, subtask regeneration, or replanning.
Table 3: Automated metric results on AVE-Compass, grouped into Cross-Modal, Video, and Audio metrics. All metrics are higher-is-better. Best in bold. ∗Gemini misses 16 speech edits due to content moderation. †Speech Quality and Lip Sync are computed only on speech-category edits.
ModelCross-ModalVideoAudio
Lip Sync†AV SyncVideo AestheticSubject ConsistencyMotion SmoothnessAudio AestheticSpeech Quality†
AVE-Agent (Wan)0.622+0.0910.766+0.0730.452+0.0010.972+0.0030.987+0.0010.614+0.0050.368-0.060
Wan2.70.5310.6930.4510.9690.9860.6090.428
HappyHorse0.6200.6950.4390.9750.9890.6500.726
Gemini-Omni∗0.7010.4340.9740.9880.627
Seedance0.4310.7180.4300.9710.9870.6290.388
LTX20.6180.7580.4680.9680.9860.6410.522
Figure 4: Robustness under difficulty. Left: Editing Intent across source-video duration. Right: object-localization hardness, audio-source complexity, and cross-modal linkage degree evaluated with modality-matched metrics. Hatched audio-source complexity bars mark Audio Fidelity Preserving inflated by audio non-response. Scores are on a 0–100 scale.
Figure 4: Robustness under difficulty. Left: Editing Intent across source-video duration. Right: object-localization hardness, audio-source complexity, and cross-modal linkage degree evaluated with modality-matched metrics. Hatched audio-source complexity bars mark Audio Fidelity Preserving inflated by audio non-response. Scores are on a 0–100 scale.
Table 4: Response-gated Fidelity Preserving and Realism by modality. Response rate measures whether the target modality changes. The “all” columns average over all evaluated cases, while the “gated” columns recompute Fidelity Preserving and Realism only on cases where the corresponding target modality responds. All metrics are higher-is-better. Bold marks the lowest, i.e., worst, response rates.
ModelResponse RateFidelity Preserving (all)Fidelity Preserving (gated)Realism (all)Realism (gated)
VideoAudioVideoAudioVideoAudioVideoAudioVideoAudio
AVE-Agent (Wan)92.794.480.174.877.970.945.378.943.076.6
Wan2.790.695.678.548.176.344.843.577.440.475.1
HappyHorse95.874.474.553.373.131.749.976.048.873.2
Gemini-Omni89.722.076.896.174.280.649.284.546.881.2
Seedance69.147.883.080.474.960.154.084.448.582.8
LTX290.191.124.542.316.839.042.086.236.784.4
Figure 5: Metric orthogonality, single-modality editing, and error analysis. Left: Spearman correlation between subjective and objective metrics. Middle: video-only and audio-only performance, with axes normalized per metric. Right: per-model error counts across five failure categories.
Figure 5: Metric orthogonality, single-modality editing, and error analysis. Left: Spearman correlation between subjective and objective metrics. Middle: video-only and audio-only performance, with axes normalized per metric. Right: per-model error counts across five failure categories.
Table 5: Human-LLM agreement of MLLM-as-Judge metrics. Agreement measures consistency between human annotations and automatic MLLM judgments, averaged over evaluated models. Values are percentages.
MetricEditing IntentInstruction FollowingFidelity PreservingRealism
OverallVideoAudioOverallVideoAudioOverallVideoAudioOverallVideoAudio
Agreement89.990.089.991.392.989.788.887.990.191.093.089.0
Figure 6: Ablation of planner-side prompt enhancement and evaluator-driven retry refinement. Solid bars show the score without the module; hatched segments show the drop from the full AVE-Agent.
Figure 6: Ablation of planner-side prompt enhancement and evaluator-driven retry refinement. Solid bars show the score without the module; hatched segments show the drop from the full AVE-Agent.
Table 6: Reference-free objective score routing in AVE-Compass. Each score is activated only for the instruction categories where it provides relevant evidence.
Metric GroupScore NameActive Categories
Cross-Modallip_syncspeech
av_syncaudio-only, joint, speech
Videovideo_aestheticvideo-only, joint, speech
subject_consistencyvideo-only, joint, speech
motion_smoothnessvideo-only, joint, speech
Audioaudio_aestheticaudio-only, joint, speech
speech_qualityspeech
Figure 7: Qualitative case analysis. Each row shows the source frame, edit instruction, and aligned edited frames from AVE-Agent, LTX2, Wan2.7, Seedance, Gemini-Omni, and HappyHorse. Red tags mark the dominant failure observed in each selected output, including low Fidelity Preserving under regeneration, over-realistic visual effects, non-response, physical-logic errors, spurious text/object insertion, temporal inconsistency, and poor Audio Fidelity Preserving.
Figure 7: Qualitative case analysis. Each row shows the source frame, edit instruction, and aligned edited frames from AVE-Agent, LTX2, Wan2.7, Seedance, Gemini-Omni, and HappyHorse. Red tags mark the dominant failure observed in each selected output, including low Fidelity Preserving under regeneration, over-realistic visual effects, non-response, physical-logic errors, spurious text/object insertion, temporal inconsistency, and poor Audio Fidelity Preserving.
Table 7: Difficulty-stratified analysis on AVE-Compass, each axis paired with its modality-matched metric. All values are on a 0–100 scale. Video Realism is normalized from the 1–5 rubric via (x−1)/4×100. Baselines degrade with difficulty, while our agent stays notably more robust under audio-source complexity and implicit cross-modal linkage.
ModelObject-Localization Hardness Video-IFObject-Localization Hardness Video RealismAudio-Source Complexity Audio Fidelity PreservingCross-Modal Linkage Audio-IF
easyhardeasyhardsimplemod.complexexplicitimplicit
AVE-Agent (Wan)90.583.748.544.773.675.175.969.073.3
Wan2.789.376.347.342.760.346.336.061.356.0
HappyHorse78.674.956.648.566.255.038.061.313.3
Gemini-Omni73.174.254.448.1100.096.491.410.513.3
Seedance64.348.159.852.780.585.576.821.639.3
LTX275.672.237.343.052.035.836.465.472.7
Table 8: Single-modality editing (objective metrics). Top: video-only (n=16)—visual quality plus Audio Similarity (preservation of the untouched audio). Bottom: audio-only (n=5)—audio quality plus Video Similarity (preservation of the untouched video). Best in bold. ‡Our agent’s Audio Similarity is mildly depressed by a residual ∼1-frame muxing offset (frame-level alignment leaves an ∼8 ms residual); sample-aligned, the audio is essentially preserved.
ModelVideo AestheticSubject ConsistencyMotion SmoothnessAudio Similarity
AVE-Agent (Wan)0.4630.9830.9850.969‡
Wan2.70.4670.9830.9860.759
HappyHorse0.4540.9860.9890.999
Gemini-Omni0.4380.9830.9870.978
Seedance0.4420.9800.9850.833
LTX20.4780.9730.9830.647
Table 9: Error analysis across five failure categories (error counts; lower is better). An edit is counted as an error in a category when it fails the corresponding checklist questions (per-case yes-rate <0.5) or, for audio-visual quality, when its Realism is below 3/5. Best (lowest) per row in bold; Gemini’s low Audio Fidelity Preserving count reflects non-response, not preservation quality.
CategoryAVE-Agent (Wan)Wan2.7Happy HorseGemini OmniSeedanceLTX2
1. Poor Video Fidelity Preserving (FP×video)2227342828161
2. Poor Audio Fidelity Preserving (FP×audio)379793534106
3. Poor Video Instruction Following (IF×video)203239408844
4. Poor Audio Instruction Following (IF×audio)48697814413356
5. Poor AV Quality (Realism <3/5)485850313139
Total175283294248314406
Table 10: MLLM-as-Judge agent-backbone ablation on the 34-case subset. Scores are on a 0–100 scale. All metrics are higher-is-better. Best results within each backbone pair are in bold.
ModelEditing IntentInstruction FollowingFidelity PreservingRealism
OverallVideoAudioOverallVideoAudioOverallVideoAudioOverallVideoAudio
AVE-Agent (Wan)62.5+17.169.4+10.757.0+30.581.8+7.785.3+2.079.8+14.177.1+12.280.9+8.173.5+22.858.0+0.137.7+1.778.3-1.5
Wan45.458.726.574.183.365.764.972.850.757.936.079.8
AVE-Agent (sd2)49.0+16.244.2+10.552.0+25.170.0+26.369.6+19.168.5+31.970.0-12.664.7-17.577.2-6.164.8-8.450.2-10.180.3-5.7
Seedance32.833.726.943.750.536.682.682.283.373.260.386.0
Table 11: Automated metric agent-backbone ablation on the 34-case subset. All metrics are higher-is-better. Best results within each backbone pair are in bold. †Speech Quality and Lip Sync are computed only on the two speech-category edits.
ModelCross-ModalVideoAudio
Lip Sync†AV SyncVideo AestheticSubject ConsistencyMotion SmoothnessAudio AestheticSpeech Quality†
AVE-Agent (Wan)0.858+0.0600.821+0.1090.452+0.0000.971+0.0040.986+0.0000.627+0.0310.402-0.144
Wan0.7980.7120.4520.9670.9860.5960.546
AVE-Agent (sd2)1.000+0.6710.691+0.0590.419-0.0150.976+0.0070.989+0.0010.608-0.0260.506+0.010
Seedance0.3290.6320.4340.9690.9880.6340.496

研究结果

  • AVE-Agent在编辑意图和指令遵循两项得分最高,尤其在音频侧指令遵循和音画同步方面领先幅度较大。
  • LTX2能完成许多指令但基本上重新生成了整段视频,导致保真保留分数很差;Gemini-Omni和Seedance有时直接原样返回音频或视频,这让保留分数虚高但并未满足编辑要求。
  • 部分模型在自动化的视频/音频美学指标上得分较高,但在独立的真实感评分上明显偏低,说明技术质量和物理逻辑合理性是两件不同的事。
  • 五次重复评分中,自动化指标波动小于0.01,多模态大模型评分波动小于1%,与人工标注的平均一致率在四个维度上都接近90%。
  • 错误分析显示AVE-Agent总错误数最少(175起),远低于其他基线的248至406起,且没有单一主导的失败类型;而Gemini因为几乎不改音频导致音频指令遵循错误达144起,LTX2因过度重新生成导致视频保真保留错误达161起。

可应用场景

  • 在部署前用AVE-Compass筛查需要同步修改画面和声音的音画编辑模型
  • 诊断某个模型典型的失败模式(例如过度重新生成还是不响应),指导编辑流水线的改进方向
  • 为教育、内容再配音、本地化等无障碍相关的音画编辑工作流参考这套“规划-执行-评估”智能体设计

局限与待验证事项

  • 由于调用闭源API和多模态大模型评审的单例成本较高,基准规模被有意控制在145段视频、196条指令。
  • AVE-Agent依赖第三方工具,这些工具的行为可能随时间变化,影响长期可复现性。
  • Gemini因内容审核政策未能完成16条语音相关编辑,这些案例被排除在其相关结果之外。
  • AVE-Agent的音频相似度自动指标因约1帧(约8毫秒)的残余对齐偏移而略偏低,作者说明若按样本级对齐,音频基本被完整保留。
  • 作者提示了深度伪造、非自愿声音克隆等双重用途风险,发布版本不含身份克隆工具,并建议配合水印等溯源技术使用。

为什么重要

评测或部署音画编辑模型的人往往只看画面就下判断,而这项研究用数据说明声音侧的失误其实更突出,并提供了一套具体可用的诊断工具。它也给出了挑选或改进此类编辑系统时应重点检查的失败类型清单。

本文术语

  • MLLM-as-Judge · 用大型多模态语言模型像人类评审一样给编辑结果打分的方法
  • 检查清单(checklist) · 一组是/否问题,用来核实编辑是否执行、原始内容是否被保留
  • 编辑意图(Editing Intent, EI) · 由指令遵循乘以保真保留计算得出的分数,衡量编辑是否完整且干净地完成
  • 子任务依赖图(DAG) · 表示一条指令拆分出的各个子任务之间执行顺序和跨模态关系的图结构
  • 保真保留(Fidelity Preserving, FP) · 衡量未被要求编辑的画面和声音内容是否与原始视频保持一致

论文原文摘要(英文)

While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modality often requires coordinated changes in the other. Existing benchmarks primarily evaluate visual transformations on silent clips or isolated audio editing, leaving complex audio-visual editing and cross-modal consistency underexplored. We introduce AVE-Compass, a comprehensive benchmark with 145 curated source videos, 196 audio-visually coupled ed

作者 · Yuqing Wen

在 arXiv 阅读

最新论文

全部论文 →

METAL MEDIA 最新报道

图片来源: Yuqing Wen et al., arXiv:2607.24821, CC BY 4.0