K-文化的一切——从回归到 K-美妆,发送到您的邮箱订阅邮件

METAL MEDIA

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

arXiv:2607.286092026-07-29

研究发现,用来判断AI操作电脑代理是否真正完成任务的AI裁判其实很容易被骗

能看屏幕、点击输入来操作电脑的AI代理(CUA)完成任务后,需要有人核实它是否真的完成了,这个核实工作越来越多地交给另一个AI——视觉语言模型(VLM)当裁判,但此前没人系统检验过这个裁判是否靠谱。研究团队从零搭建了覆盖网页、手机、Ubuntu、Windows四个平台的真实环境,收集并由人工标注出1019条可信的任务记录,组成OSReward基准,测试了27个VLM裁判,发现即便最强模型在难例集上准确率也只掉到七成左右,且大多裁判都有把失败任务误判为成功的宽松偏差。为此团队发布了10万样本的开放训练语料OS-Shepherd-100K,并训练出开源奖励模型OS-Shepherd(9B和35B),以商用前沿模型30到60倍更低的成本达到接近的判断准确度。

METAL MEDIA 解读图

从OSReward到OS-Shepherd的研究流程

证据状态已报告实测结果

  1. 1. 搭建真实环境在网页、手机、Ubuntu、Windows上配置真实应用、已登录账号、真实文件和干扰内容
  2. 2. 运行代理并采集人工黄金标签四个模型家族的代理执行经验证的指令,三名标注员独立打标,分歧升级给资深评审,最终形成1019条黄金任务记录(完整集/Hard/Multi)
  3. 3. 测试27个VLM裁判前沿模型到小型开源模型均在黄金标准上测试,暴露出宽松偏差和难例上的表现崩溃
  4. 4. 构建OS-Shepherd-100K语料超过30万条裁判实例被提炼成10万样本、附带推理过程的训练语料
  5. 5. 训练并验证OS-Shepherd监督微调加GRPO强化学习训练出9B/35B奖励模型,以30至60倍更低成本接近商用裁判的准确度
这是 METAL MEDIA 制作的解读图,并非论文作者提供的原图。

他们做了什么

  1. 直接复用现有基准留下的任务记录会混入质量问题和不准确的标签,因此研究团队从零搭建了网页、手机、Ubuntu、Windows四个平台的专用采集环境,配备真实应用、已登录账号、真实文件和干扰内容
  2. 来自Claude、Gemini、Kimi、Qwen四个模型家族的代理执行经人工验证的指令,产生了1019条真实成功与失败混合的任务记录,每条都由三名独立标注员打标,意见不一致时升级给两名资深评审做最终裁定,整个过程耗费约800人工小时
  3. 这套黄金标准数据被拆分为完整的OSReward集、专门收集难例的OSReward-Hard子集(284条,重新调整为30%成功/70%失败)、以及对440条成功任务打分对齐度与效率的OSReward-Multi子集,并在统一协议下测试了27个VLM裁判
  4. 表现最好的裁判(Claude-Opus-4-8)在完整集上准确率为89.7%,但在OSReward-Hard难例集上骤降至69.7%,平均裁判准确率跌到52%,且三分之二的错误都是把未完成的任务误判为成功
  5. 为弥补这一缺口,团队将超过30万条裁判实例提炼成10万样本的语料库OS-Shepherd-100K,并分两阶段(监督微调加强化学习)训练出OS-Shepherd模型,以30到60倍更低的成本达到接近商用前沿模型的判断准确度
Figure 1: Cost against binary accuracy on OSReward-Hard: reliable judges are expensive, and the OS-Shepherd models come nearest their accuracy at a fraction of the cost.
Figure 1: Cost against binary accuracy on OSReward-Hard: reliable judges are expensive, and the OS-Shepherd models come nearest their accuracy at a fraction of the cost.
Table 1: Main-setting results for the reference judges and OS-Shepherd on OSReward and OSReward-Hard along with their access status, sorted by full-set accuracy.
JudgeAccessOSRewardOSReward-Hard
AccsRecfRecBalAccAccsRecfRecBalAcc
Claude-Opus-4-8closed89.791.188.990.069.769.869.769.7
GPT-5.5closed89.591.887.889.867.366.367.767.0
Claude-Opus-4-6closed89.592.787.790.267.372.165.268.6
Gemini-3.1-Proclosed87.990.286.288.261.661.661.661.6
Gemini-3.5-Flashclosed87.895.781.888.859.581.450.065.7
Claude-Sonnet-4-6closed87.797.580.388.959.290.745.568.1
GPT-5closed87.486.887.987.458.143.064.653.8
GPT-5.4closed87.187.387.087.163.062.863.163.0
Gemini-3-Flashclosed87.096.679.888.257.086.044.465.2
GPT-5-miniclosed86.193.880.287.056.379.146.562.8
Kimi-K2.5open weights85.995.579.287.354.883.742.162.9
Qwen3.5-397B-A17Bopen weights85.895.278.686.958.591.943.967.9
GPT-5.4-miniclosed85.282.587.284.958.148.262.455.3
Claude-Haiku-4-5closed84.580.987.284.059.547.764.656.2
GPT-5.2closed83.973.092.282.663.030.277.353.8
Gemini-2.5-Flashclosed83.395.574.084.848.990.730.860.8
Doubao-2.0-Liteclosed83.398.572.185.345.596.124.360.2
GPT-5-nanoclosed82.397.071.184.145.495.323.759.5
Intern-S1-Proopen weights82.392.374.783.543.770.931.851.4
Qwen3.5-35B-A3Bopen weights82.292.474.583.551.183.736.960.3
Qwen3.5-27Bopen weights82.097.470.584.044.292.923.258.0
GPT-4oclosed81.096.869.082.939.490.717.253.9
Intern-S2-Previewopen weights80.698.466.982.740.394.216.855.5
Qwen3.5-122B-A10Bopen weights79.696.866.481.639.489.517.753.6
Qwen3-VL-8Bopen weights77.199.859.979.836.2100.08.254.1
Qwen3-VL-235Bopen weights74.099.154.977.031.497.72.550.1
Qwen3-VL-30Bopen weights69.499.846.373.031.198.81.550.2
OS-Shepherd-9B (ours)open weights + data86.186.686.086.360.266.357.661.9
OS-Shepherd-35B-A3B (ours)open weights + data85.685.086.285.662.768.660.164.3
Figure 2: From realistic environments to raw trajectories: annotators prepare the environments and write grounded instructions on them, agents from four model families execute the instructions.
Figure 2: From realistic environments to raw trajectories: annotators prepare the environments and write grounded instructions on them, agents from four model families execute the instructions.
Table 2: Strong judges on OSReward-Multi (%), sorted by AUC; best per column in bold.
JudgeMacro-recallAUC
AlignEfficMulti
GPT-5.558.768.263.566.7
Claude-Opus-4-852.968.760.865.6
Claude-Sonnet-4-653.262.657.961.9
Gemini-3.5-Flash47.671.459.560.8
OS-Shepherd-35B-A3B (ours)47.765.856.860.7
OS-Shepherd-9B (ours)44.154.049.058.5
Gemini-3-Flash50.661.556.055.8
Figure 3: The annotation pipeline. Each pre-filtered trajectory is labeled by three independent annotators; disagreements go to meta review, and the verified gold set is read as three views
Figure 3: The annotation pipeline. Each pre-filtered trajectory is labeled by three independent annotators; disagreements go to meta review, and the verified gold set is read as three views
Table 3: OS-Shepherd-100K judge-instance pool by platform.
PlatformInstancesShare
Web119,46937%
Windows62,05319%
macOS45,02814%
Ubuntu (GUI only)34,35511%
Ubuntu (GUI + CLI)29,7859%
Mobile30,94110%
Total321,631100%
Figure 4: OSReward at a glance: outcome composition of the full and Hard sets, platform mix, and trajectory length; failed runs are markedly longer.
Figure 4: OSReward at a glance: outcome composition of the full and Hard sets, platform mix, and trajectory length; failed runs are markedly longer.
Table 4: OS-Shepherd against its untuned base, on the full set and OSReward-Hard.
ModelOSRewardOSReward-Hard
AccsRecfRecBalAccAccsRecfRecBalAcc
Qwen3.5-9B (base)76.798.959.979.439.497.714.155.9
OS-Shepherd-9B86.186.686.086.360.266.357.661.9
Qwen3.5-35B-A3B (base)82.292.474.583.551.183.736.960.3
OS-Shepherd-35B-A3B85.685.086.285.662.768.660.164.3
Figure 5: Judge bias on OSReward (left) and OSReward-Hard (right). Judges below the diagonal skew lenient; most of the field sits there, and the skew widens on the hard set.
Figure 5: Judge bias on OSReward (left) and OSReward-Hard (right). Judges below the diagonal skew lenient; most of the field sits there, and the skew widens on the hard set.
Table 5: Action spaces of the executing agents on web (left) and mobile (right). On web, the second block lists the browser primitives and the third the terminal action.
ActionDescription
click [coord]Clicks at the specified screen location.
double_click [coord]Double-clicks at the specified screen location.
hover [coord]Moves the pointer to the specified screen location.
scroll [up/down]Scrolls the screen in the specified direction.
drag [coord] [coord]Drags from the first coordinate to the second.
type [text]Types text at the current cursor location.
fill [coord] [text]Clicks at a location, clears its content, and types text.
clear [coord]Clicks at a location and clears the current text input.
hotkey [keys]Presses the specified key or key combination.
wait [seconds]Waits for the page to load.
goto [url]Navigates directly to a URL.
go_backNavigates to the previous page in browser history.
go_forwardNavigates to the next page in browser history.
select_option [coord] [text]Selects text from the dropdown at a screen location.
set_checked [coord] [bool]Sets the control state at a screen location.
stop [answer]Terminates the episode and returns the final answer.
Figure 6: Per-judge error composition, seven representative judges: over-accepting an incomplete task (warm colors) dominates every family.
Figure 6: Per-judge error composition, seven representative judges: over-accepting an incomplete task (warm colors) dominates every family.
Table 6: Action space of the executing agents on Windows, with each action’s parameter format.
ActionParameter specification
computer.mouse.move_absFormat: [x,y] Details: Move the mouse to a normalized screen position; x, y are floats.
computer.mouse.single_clickFormat: [] Details: Single-click at the current mouse position.
computer.mouse.double_clickFormat: [] Details: Double-click at the current mouse position.
computer.mouse.right_clickFormat: [] Details: Right-click at the current mouse position.
computer.mouse.scrollFormat: [direction] Details: Scroll the screen up or down; direction is a string.
computer.mouse.dragFormat: [x1,y1,x2,y2] Details: Drag from the current mouse position to the target normalized position; coordinates are floats.
computer.keyboard.writeFormat: [text] Details: Type the given text.
computer.keyboard.pressFormat: [key] Details: Press a keyboard key such as Enter or Delete.
computer.os.open_programFormat: [program_name] Details: Open the specified application.
computer.window_manager.switch_to_applicationFormat: [window_name] Details: Switch to the specified open window or application.
computer.waitFormat: [time] Details: Wait for the given number of milliseconds (time is an integer).
COMMANDFormat: [] Details: Output and execute a Python code block for the current step.
ANSWERFormat: [answer] Details: Return the specific answer text for the given prompt.
DONEFormat: [] Details: The task is finished; end the episode.
FAILFormat: [] Details: The task cannot be completed; end the episode.
Figure 7: Mean per-judge binary accuracy on OSReward-Hard, by platform and by failure type. Failure types are multi-label, so their counts sum to more than the number of failed trajectories.
Figure 7: Mean per-judge binary accuracy on OSReward-Hard, by platform and by failure type. Failure types are multi-label, so their counts sum to more than the number of failed trajectories.
Table 7: Application coverage of the collection infrastructure on Ubuntu, Windows, and Android, grouped by function; applications marked with ∗ require a signed-in account. The web platform runs on live websites rather than installed applications (Section A.1).
PlatformGroupApplications
UbuntuWeb & communicationChrome, Thunderbird, Zoom∗
DevelopmentVS Code, PyCharm, GitKraken, DBeaver, Wireshark, Meld, terminal
DocumentsLibreOffice Writer / Calc / Impress, TeXstudio, PDF Arranger, Zotero, Calendar
Graphics & designGIMP, Blender, Inkscape, Krita, Darktable, LibreCAD, KiCad, draw.io
MediaVLC, Audacity, Mixxx, HandBrake, Shotcut, OBS Studio, MuseScore, Spotify∗
ScientificScilab, KAlgebra, GRASS GIS, Google Earth Pro, ChimeraX, Celestia
PersonalHomeBank
WindowsWeb & communicationChrome, Microsoft Edge, Thunderbird, Feishu∗, Discord∗, Zoom∗, Tencent Meeting∗
DevelopmentVS Code, PyCharm, DBeaver
DocumentsNotepad, PDF Arranger, Zotero
Graphics & designBlender, Krita, draw.io
MediaVLC, Shotcut, HandBrake, Spotify
UtilitiesFile Explorer, Calculator
PersonalSteam∗
AndroidWeb & communicationBrowser, Firefox, Gmail∗, SMS, Contacts
DocumentsMarkor, Google Keep∗, Calendar
Graphics & designDraw
MediaCamera, Gallery, Google Photos∗, Audio Recorder, VLC, Retro Music
Maps & navigationOsmAnd, Google Maps
PersonalExpense, Recipe, Yahoo Finance
UtilitiesFiles, Clock, Calculator, and system tasks
Figure 8: Overview of input ablations: Δ binary accuracy per (setting × model) vs. the main setting.
Figure 8: Overview of input ablations: Δ binary accuracy per (setting × model) vs. the main setting.
Table 8: Action space of the executing agents on Ubuntu, with each action’s parameter format.
ActionParameter specification
clickFormat: [desc,num_clicks,button,hold_keys] Details: Target element description; clicks number; button to click; keys to hold.
typeFormat: [desc,text,overwrite,enter,terminal] Details: Target element description; text content; overwrite flag (bool); press enter after typing (bool); terminal flag (bool).
scrollFormat: [desc,clicks,shift] Details: Target element description; clicks (+up/−down); shift for horizontal scroll (bool).
drag_and_dropFormat: [start_desc,end_desc,hold_keys] Details: Descriptions for start/end locations; keys to hold during drag.
hotkeyFormat: [keys] Details: List of keys to press in combination (e.g., [‘ctrl’, ‘c’]).
hold_and_pressFormat: [hold_keys,press_keys] Details: Keys to hold down while pressing a sequence of other keys.
openFormat: [app_or_filename] Details: Name of the application or file to open.
call_code_agentFormat: [task] Details: A self-contained goal executable via code (e.g., data analysis, file processing).
waitFormat: [time] Details: Time to wait in seconds.
doneFormat: [] Details: Signals successful completion of the entire task.
failFormat: [] Details: Signals that the task is impossible to complete.
Figure 9: The OS-Shepherd-100K pipeline: self-collected and open-source trajectories are ensemble-judged and distilled into the training set. Band widths ∝ trajectory counts.
Figure 9: The OS-Shepherd-100K pipeline: self-collected and open-source trajectories are ensemble-judged and distilled into the training set. Band widths ∝ trajectory counts.
Table 9: All evaluated models: the 27 reference judges (top, by full-set accuracy) and our reward models. The last column lists the extra thinking or reasoning-effort levels beyond the main setting; access classes are in Table 1.
JudgeAPI identifierThinking levels
Claude-Opus-4-8 (Anthropic 2026b)claude-opus-4-8
GPT-5.5 (OpenAI 2026b)gpt-5.5medium/high/xhigh
Claude-Opus-4-6 (Anthropic 2026a)claude-opus-4-6xhigh/max
Gemini-3.1-Pro (Gemini Team 2025)gemini-3.1-pro-preview
Gemini-3.5-Flash (Gemini Team 2025)gemini-3.5-flash
Claude-Sonnet-4-6 (Anthropic 2026c)claude-sonnet-4-6xhigh/max
GPT-5 (OpenAI 2025b)gpt-5
GPT-5.4 (OpenAI 2026a)gpt-5.4
Gemini-3-Flash (Gemini Team 2025)gemini-3-flash-preview
GPT-5-mini (OpenAI 2025b)gpt-5-mini
Kimi-K2.5 (Kimi Team et al. 2026)kimi-k2.5
Qwen3.5-397B-A17B (Qwen Team 2026)qwen3.5-397b-a17btwo settings
GPT-5.4-mini (OpenAI 2026a)gpt-5.4-mini
Claude-Haiku-4-5 (Anthropic 2025)claude-haiku-4-5-20251001
GPT-5.2 (OpenAI 2025b)gpt-5.2
Gemini-2.5-Flash (Comanici et al. 2025)gemini-2.5-flash
Doubao-2.0-Lite (ByteDance Seed Team 2026)doubao-seed-2-0-lite-260428
GPT-5-nano (OpenAI 2025b)gpt-5-nano
Intern-S1-Pro (Zou et al. 2026)intern-s1-pro
Qwen3.5-35B-A3B (Qwen Team 2026)qwen3.5-35b-a3b
Qwen3.5-27B (Qwen Team 2026)qwen3.5-27b
GPT-4o (Hurst et al. 2024)gpt-4o
Intern-S2-Preview (Zou et al. 2026)intern-s2-preview
Qwen3.5-122B-A10B (Qwen Team 2026)qwen3.5-122b-a10b
Qwen3-VL-8B (Bai et al. 2025)qwen3-vl-8b-instructtwo settings
Qwen3-VL-235B (Bai et al. 2025)qwen3-vl-235b-a22b-instruct
Qwen3-VL-30B (Bai et al. 2025)qwen3-vl-30b-a3b-instruct
OS-Shepherd-9B (ours)os-shepherd-9b
OS-Shepherd-35B-A3B (ours)os-shepherd-35b-a3b
Figure 10: Judges on three existing CUA benchmarks against each benchmark’s human-written verifier (matched subsets): accuracy and failure recall; means are computed before rounding.
Figure 10: Judges on three existing CUA benchmarks against each benchmark’s human-written verifier (matched subsets): accuracy and failure recall; means are computed before rounding.
Table 10: OSReward beside existing CUA reward works, on data provenance and released artifacts rather than head-to-head scores (their input formats and platform scopes preclude a shared protocol). OSReward is the only one built end-to-end from freshly collected, human-gold trajectories and the only one whose gold goes beyond a binary verdict. Platforms W/M/D = web/mobile/desktop; Instr. / Traj. / Gold flag fresh instructions, fresh trajectories, and human-labeled gold; Corpus / Model give any released training corpus and reward model (✓ yes, ✗ no, ∼ partial, – n/a).
Reward benchmarkReward model
DatasetPlatformsActionInstr.Traj.GoldLabelsCorpusModel
OSReward (ours)W, M, DGUI+CLIBinary + fine-grained✓ 100K✓ 9B/35B
Web-Shepherd Chae et al. 2025WGUIChecklist✓ 40K✓ 3B/8B
GUI-Shepherd Chen et al. 2025aMGUI✓ 52K✓ 7B
CUARewardBench Lin et al. 2025DGUIBinary
OS-Themis Li et al. 2026W, M, DGUIBinary
Figure 11: The twenty most frequent task types among the roughly 32K web instructions sampled for collection, together about 65% of the pool. Fact-finding lookups dominate, followed by article, academic-paper, and shopping tasks.
Figure 11: The twenty most frequent task types among the roughly 32K web instructions sampled for collection, together about 65% of the pool. Fact-finding lookups dominate, followed by article, academic-paper, and shopping tasks.
Table 11: The OS-Shepherd-100K judge-instance pool by source (321,631 instances over eight sources). success is the share of agent-successful verdicts per source; the web pool is the most failure-rich. Nothing is drawn from any existing benchmark’s test set (Sections A.1 and 7).
SourcePlatformInstancessuccess
Self-collectedWebWeb117,25145%
Ubuntu (GUI+CLI)Ubuntu29,78572%
Scientific (Sun et al. 2026b)Ubuntu14,33959%
WindowsWindows3,59950%
OS-Genesis (re-generated; Sun et al. 2025a)Web2,21873%
ReusedOpenCUA (Wang et al. 2025)Windows / macOS103,48269%
OpenMobile (Cheng et al. 2026)Mobile30,94162%
OpenCUA (Wang et al. 2025)Ubuntu18,91678%
ScaleCUA (Liu et al. 2026)Ubuntu1,10064%
Figure 12: Failure-type profile over OSReward’s fail trajectories (multi-label shares; the catch-all others tag is excluded). A single run can carry several tags.
Figure 12: Failure-type profile over OSReward’s fail trajectories (multi-label shares; the catch-all others tag is excluded). A single run can carry several tags.
Table 12: Screenshot-setting mix of the retained training samples.
Screenshot settingShare
Last-5 frames45.1%
First-1 + last-226.0%
Last-3 frames18.9%
Last-10 frames8.9%
Last-6/7/8 frames1.2%
Table 13: OS-Shepherd training configuration for both sizes. SFT is shared (same corpus and schedule); the two RL runs share the mined set and differ only in the base checkpoint.
SFTRL
(both sizes)9B35B-A3B
Base modelQwen3.5-9B / Qwen3.5-35B-A3B9B SFT ckpt35B SFT ckpt
Samples96.6K3.1K (shared)
Rollouts / sample8 (at T=1.0, top-p 1.0)
Batch size16
Learning rate1​e−6
KL to SFT ref.0.001 (low-variance, as loss)
Max prompt / resp.24,576 / 512 tokens
Steps1 epoch∼150 (≈1 pass)
Frameworkverl + SGLang rollout back-end
Hardware32× NVIDIA H200 (4 nodes × 8)
Table 14: OS-Shepherd-9B beside its full-set accuracy tier and two frontier judges. Cost is list price to judge the full set; full/hard are binary accuracy (%).
JudgeWeightsCost ($)FullHard
Claude-Opus-4-8closed86.0489.769.7
GPT-5.5closed45.4489.567.3
Kimi-K2.5open20.3785.954.8
Qwen3.5-397B-A17Bopen7.9685.858.5
GPT-5.4-miniclosed6.2085.258.1
GPT-5-miniclosed2.1786.156.3
Gemini-3-Flashclosed2.0287.057.0
OS-Shepherd-9B (ours)open1.3686.160.2
Qwen3.5-9Bopen1.3676.739.4
Table 15: Thinking and reasoning effort. Each left-hand row contrasts two settings of one model, so Δ is within-model; the Qwen3-VL-8B thinking arm rejects ∼6%, making its Δ intersection-paired. Right: the GPT-5.5 reasoning-effort sweep.
ModelSettingAccSettingAccΔ
Qwen3-VL-8Bno thinking77.1thinking81.7+2.83
Qwen3.5-397B-A17Bno thinking85.8thinking86.7+0.89
Claude-Sonnet-4-6xhigh87.7max88.5+0.59
Claude-Opus-4-6xhigh89.5max90.0+0.39

研究结果

  • 最佳裁判Claude-Opus-4-8在完整OSReward集上准确率为89.7%,但在难例集OSReward-Hard上,所有裁判的准确率都下降了20到43个百分点,最好的模型也只达到69.7%
  • 约三分之二的裁判错误是把未完成的任务误判为成功,这是所有被测裁判共同的主要错误类型,每个模型至少有48%的错误属于这一类
  • OS-Shepherd-9B和35B在相同评测条件下,以30到60倍更低的成本达到接近商用前沿裁判的准确度,强化学习阶段将验证准确率从约70%提升到约77%
  • 在外部基准OSWorld上,88%的裁判错误是误报成功,超过16步的长任务记录中,准确率从0.76降到0.57,误报率从0.20升到0.37

可应用场景

  • 用OS-Shepherd作为低成本奖励信号,在强化学习训练中大批量给代理的任务记录打分
  • 在数据采集流程中用于人工审核前的初筛,批量判断大量代理运行记录是否成功
  • 在正式采用某个新VLM作为CUA裁判前,用OSReward-Hard这类难例集检验其是否存在宽松偏差

局限与待验证事项

  • 该基准只覆盖网页、手机、Ubuntu和Windows,不包含macOS(训练语料中的macOS部分来自另一份复用的公开数据集)
  • 评测协议默认只用最后五张截图这一固定设置进行测试,在其他输入方式(如允许工具调用、逐步监督)下的可靠性尚未在本研究中验证
  • 论文提出,明确要求裁判核实任务是否真正完成的提示方式,可能比论文中测试的集成方法更能缓解宽松偏差,但这一点留给未来研究
  • OS-Shepherd的强化学习数据因虚假成功案例主要集中在桌面平台,导致移动端和网页端样本相对不足

为什么重要

训练和评估AI代理越来越依赖另一个AI来判断每次任务是否真正完成,如果这个裁判本身不可靠却不被察觉,训练数据和评测分数都可能被系统性地扭曲。这项研究首次以标准化方式测出这种可靠性的真实水平,并提供了一个开源、低成本的替代方案,对所有在做电脑操作代理开发或评测的人都有直接意义。

本文术语

  • CUA(操作电脑的AI代理) · 通过看屏幕、点击、输入等方式直接操作电脑完成任务的AI代理
  • VLM(视觉语言模型) · 能同时理解图像(如屏幕截图)和文字的AI模型
  • 虚假成功(false success) · 代理声称任务已完成,但实际上并未达成目标的情况
  • 奖励模型 · 给另一个AI的输出打分,用于指导强化学习训练或筛选数据的模型
  • GRPO · OS-Shepherd强化学习阶段所用的策略优化算法

论文原文摘要(英文)

Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the task instruction is central to CUA evaluation, data curation, and reinforcement learning. Neither human-written verifiers nor human annotators can provide such verification at scale, so the field increasingly turns to vision-language models (VLMs) as judges of CUA trajectories. But a fundamental question has long gone

作者 · Qiushi Sun

在 arXiv 阅读

最新论文

全部论文 →

METAL MEDIA 最新报道

图片来源: Qiushi Sun et al., arXiv:2607.28609, CC BY 4.0