K-文化的一切——从回归到 K-美妆,发送到您的邮箱订阅邮件

METAL MEDIA

LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents

arXiv:2608.173932026-08-17

LEGO-RL框架解决了用强化学习训练编程AI智能体时的信号失真和错位问题

LEGO-RL能在不改动Claude Code、OpenHands SDK、OpenCode等现有编程智能体程序内部逻辑的前提下,把它们接入强化学习训练流程。它同时解决了两个核心问题:环境崩溃或投机取巧导致的奖励信号被污染,以及智能体实际生成的内容与训练时重新计算的结果对不上。用LEGO-RL在三种智能体环境下训练Qwen3.5-35B-A3B模型后,在SWE-bench Verified编程基准上的解题率都有明显提升。

METAL MEDIA 解读图

LEGO-RL框架解决了用强化学习训练编程AI智能体时的信号失真和错位问题

  1. 01用强化学习训练编程智能体需要在训练时精确还原智能体当时生成的每一个词元,但现有的智能体运行框架常常会重写或压缩自己的历史记录,导致还原失败
  2. 02LEGO-RL在智能体调用大模型的接口处植入一个内部代理,直接捕获真实生成的词元和概率值;对于只激活部分子网络的混合专家模型,还会记录并在训练时重放当时具体用了哪些专家
  3. 03在隔离的沙盒环境中,系统通过镜像缓存加快启动速度,并设置分阶段防御机制来阻止智能体偷看评分信息等作弊行为,同时把失败的运行过滤掉,避免污染训练
  4. 04一个实时监控界面让研究人员能够追踪训练曲线异常的具体原因,定位到哪一步失败、哪些任务解决了、哪些没解决
  5. 05在OpenHands SDK上解题率从64.0%升到70.4%,在Claude Code上从62.4%升到68.2%,在OpenCode上从57.2%升到66.6%,同时训练时与生成时的概率一致性始终保持在0.99以上
这是 METAL MEDIA 制作的解读图,并非论文作者提供的原图。

他们做了什么

  1. 用强化学习训练编程智能体需要在训练时精确还原智能体当时生成的每一个词元,但现有的智能体运行框架常常会重写或压缩自己的历史记录,导致还原失败
  2. LEGO-RL在智能体调用大模型的接口处植入一个内部代理,直接捕获真实生成的词元和概率值;对于只激活部分子网络的混合专家模型,还会记录并在训练时重放当时具体用了哪些专家
  3. 在隔离的沙盒环境中,系统通过镜像缓存加快启动速度,并设置分阶段防御机制来阻止智能体偷看评分信息等作弊行为,同时把失败的运行过滤掉,避免污染训练
  4. 一个实时监控界面让研究人员能够追踪训练曲线异常的具体原因,定位到哪一步失败、哪些任务解决了、哪些没解决
  5. 在OpenHands SDK上解题率从64.0%升到70.4%,在Claude Code上从62.4%升到68.2%,在OpenCode上从57.2%升到66.6%,同时训练时与生成时的概率一致性始终保持在0.99以上
Figure 1: Overview of the Lego-RL training infrastructure.
Figure 1: Overview of the Lego-RL training infrastructure.
Table 1: Comparison of representative agentic RL frameworks. ✓: supported; △: partial/conditional support; –: not reported. R3: rollout routing replay. Observability denotes monitoring training runs and diagnosing execution- or trajectory-level failures.
Harness-native fidelityExecution & rewardObservability
FrameworkBlack-box harnessToken-in/ Token-outHistory alignmentR3Fully asyncSandbox executionReward-hack defenseTraining observability
verl (Sheng et al. 2025)
slime (THUDM 2025)
MOLT (NVIDIA NeMo 2026)
SkyRL-Agent (Cao et al. 2025)
AReaL (Fu et al. 2025)
Agent Lightning (Luo et al. 2025)
Polar (Xu et al. 2026)
rLLM (Berkeley Sky Computing Lab 2026)
OpenForgeRL (Yu et al. 2026)
ALE (ROLL/ROCK) (Wang et al. 2025a)
Lego-RL
Figure 2: Closed-loop operational workflow of Lego-RL. The five stages cover data preparation, run validation, training, live observability, and human review. Stage (3) corresponds to the training infrastructure shown in Figure 1, while the agent plugin acts as the control plane.
Figure 2: Closed-loop operational workflow of Lego-RL. The five stages cover data preparation, run validation, training, live observability, and human review. Stage (3) corresponds to the training infrastructure shown in Figure 1, while the agent plugin acts as the control plane.
Table 2: SWE-bench Verified performance across the three coding agents. All numbers are measured by us under the same harness version and evaluation protocol (temperature 0.7, 200 turns, 200k context budget).
Coding agentModelSWE-bench Verified (%)
OpenHands SDKQwen3.5-35B-A3B (Qwen Team 2026a)64.0
Qwen3.6-35B-A3B (Qwen Team 2026b)67.4
KAT-Coder-V2.5-Dev (KwaiKAT Team 2026)67.0
Lego-RL-Qwen3.5-35B-A3B70.4​(+6.4)
Claude CodeQwen3.5-35B-A3B (Qwen Team 2026a)62.4
Qwen3.6-35B-A3B (Qwen Team 2026b)63.4
KAT-Coder-V2.5-Dev (KwaiKAT Team 2026)66.8
Lego-RL-Qwen3.5-35B-A3B68.2​(+5.8)
OpenCodeQwen3.5-35B-A3B (Qwen Team 2026a)57.2
Qwen3.6-35B-A3B (Qwen Team 2026b)60.6
KAT-Coder-V2.5-Dev (KwaiKAT Team 2026)64.8
Lego-RL-Qwen3.5-35B-A3B66.6​(+9.4)
Figure 3: Training behavior of OpenHands SDK, Claude Code, and OpenCode over three epochs (126 training steps), showing training reward, validation reward, policy entropy, and mean response length.
Figure 3: Training behavior of OpenHands SDK, Claude Code, and OpenCode over three epochs (126 training steps), showing training reward, validation reward, policy entropy, and mean response length.
Table 3: Rollout-to-training alignment over the three matched production runs: probabilities captured at the serving boundary against trainer-side recomputation over the corresponding assistant tokens. All statistics are medians over training steps.
|Δ​log⁡p¯| per trajectory (×10−3)
Coding agentPearson rKL (×10−3)p50p90p99
OpenHands SDK0.99930.750.71.22.1
Claude Code0.99801.350.71.32.7
OpenCode0.99930.600.61.12.0
Figure 4: Trajectory termination profiles across agent scaffolds. Bars show the proportion of trajectories by termination reason; timeout and environment-setup failures are excluded from optimization.
Figure 4: Trajectory termination profiles across agent scaffolds. Bars show the proportion of trajectories by termination reason; timeout and environment-setup failures are excluded from optimization.
Table 4: Stage-wise wall-clock statistics across 3,699 OpenHands SDK training trials. Mean-time fraction is computed relative to the mean total trial duration; Other includes scheduling, trajectory handling, and sandbox cleanup.
StageMean (s)p50p90p99Mean-time fraction
Sandbox setup21.67.741.8275.22.3%
Agent setup4.24.04.88.40.5%
Agent execution840.5708.41604.32801.191.3%
Verification35.94.129.6928.53.9%
Other20.38.112.3965.52.2%
Trial total920.4770.01742.73189.1100%
Figure 5: In-batch reward distributions across training. (a–c) Distribution of tasks by the number of successful rollouts out of eight in the first and last epoch for each scaffold. The 0/8 and 8/8 groups provide no group-relative advantage. (d) Combined proportion of these zero-variation groups across epochs.
Figure 5: In-batch reward distributions across training. (a–c) Distribution of tasks by the number of successful rollouts out of eight in the first and last epoch for each scaffold. The 0/8 and 8/8 groups provide no group-relative advantage. (d) Combined proportion of these zero-variation groups across epochs.
Table 5: Ablation of sandbox optimizations.
Median latencyPaired ratio
OptimizationStageWithWithoutMedianp10–p90n
Lazy image pullsandbox setup1.57 s2.66 s1.7×100
Prebuilt task imagesandbox setup1.04 s36.2 s33.2×17.9–67.5×50
Mounted agent runtimeagent setup0.51 s7.82 s15.4×14.5–16.6×50
Packaged grading toolchainverification3.81 s2.72 s0.71×0.67–0.73×50
Figure 6: Task-selection ablation across four 951-task pools. (a) Held-out validation solve rate, measured over the validation tasks that executed. (b) Training verifier reward; levels are pool-specific, so only the slopes are comparable.
Figure 6: Task-selection ablation across four 951-task pools. (a) Held-out validation solve rate, measured over the validation tasks that executed. (b) Training verifier reward; levels are pool-specific, so only the slopes are comparable.
Table 6: Reward-integrity failure modes and their mitigations. Incidence rates are measured prior to deploying the defenses; “—” indicates cases not separately quantified.
Failure ModeIncidenceDefense
Agent-side: shortcut exploitation
Reads git history4.6–20.5%Rebase history to a single commit during agent phase; restore before grading
Downloads reference fix1.9%Per-phase egress firewall in privilege-separated sidecar
Edits test files2.4–19.4%Withhold tests until grading; revert test-path edits
Environment-side: reward detached from agent
Grader applies reference patch2.5%Audit affected instances out of pool; flag live if degenerate reward propagates
Grader requires network accessPackage all grade-time dependencies; ensure deterministic invocation
Incomplete repository buildHermetic fail-fast build; report setup failure explicitly, not as zero reward
Figure 7: Failure diagnosis with the Live UI. (a) Per-step termination reasons for an environment-failure run. (b) Assisted analysis of a collapsed run and the corresponding early-stop condition. (a) and (b) are two different diagnostic runs, neither is the Claude Code production run reported elsewhere in this section.
Figure 7: Failure diagnosis with the Live UI. (a) Per-step termination reasons for an environment-failure run. (b) Assisted analysis of a collapsed run and the corresponding early-stop condition. (a) and (b) are two different diagnostic runs, neither is the Claude Code production run reported elsewhere in this section.
Table 7: Routing-replay configurations compared. Expert overlap and top-1 agreement are undefined when replay is disabled.
Routing replay configurationPearson rMean |Δ​p|Expert overlapTop-1 agreement
Disabled0.99460.0062
Enabled, misaligned0.75030.09540.0830.026
Enabled, aligned0.99930.00250.9960.985
Figure 8: Behavioral analysis with the Live UI, Claude Code run. (a) Tool-use trajectories for eight rollouts of one task. (b) Task-level solve rates in the first and last sampled epochs.
Figure 8: Behavioral analysis with the Live UI, Claude Code run. (a) Tool-use trajectories for eight rollouts of one task. (b) Task-level solve rates in the first and last sampled epochs.
Table 8: Evidence supplied for the collapsed run. t is the least-squares slope divided by its standard error over all logged steps; the gradient norm does not clear |t|=2 and is therefore read as noise. Tokens per turn is a ratio of two rows above it and carries no separate trend statistic.
SeriesFirst five stepsLast five stepst
training reward0.3510.050−7.5
turns per trajectory18.90.96−14.8
response length (tokens)12,1702,276−7.7
tokens per turn6502,397
policy entropy0.1340.230+7.2
KL term of the loss0.00460.130+3.4
rollout–training agreement0.9950.970−3.3
gradient norm0.310.034+0.4
Figure 9: Trainer schedule under synchronous and asynchronous execution.
Figure 9: Trainer schedule under synchronous and asynchronous execution.
Table 9: Agent behaviors before and after training, over 420 trajectories at each end of the production OpenHands SDK run; the pass@k rows are over prompt groups in the first and last third of the run.
BehaviorFirstLastΔ
Reads back a file it edited73.6%98.1%+24.5
Runs the test suite85.0%93.6%+8.6
Files explored before 1st edit3.456.92+3.47
Ends with an explicit finish88.3%91.9%+3.6
Reproduces failure before editing6.7%11.2%+4.5
Solves despite a failed command63.9%66.8%+2.9
Malformed tool calls (of all calls)1.07%0.15%−0.92
Coverage, pass@883.2%87.9%+4.7
Reliability, pass828.3%39.4%+11.1
Figure 10: Lazy versus full image delivery over the same 100 task images. Nydus streams image chunks on demand, whereas OCI denotes the conventional pull, which materializes the entire image before the container starts. The panels report startup latency, cumulative network and disk traffic, and in-container read throughput.
Figure 10: Lazy versus full image delivery over the same 100 task images. Nydus streams image chunks on demand, whereas OCI denotes the conventional pull, which materializes the entire image before the container starts. The panels report startup latency, cumulative network and disk traffic, and in-container read throughput.
Table 10: Resolved hyperparameters of the three production runs (Qwen3.5-35B-A3B through the OpenHands SDK, Claude Code, and OpenCode).
Policy lossGSPO
Clip range (sequence-level)(3×10−4, 4×10−4)
Advantage estimatorGRPO
Loss aggregationseq-mean-token-mean
KL reward penaltynone
KL loss coefficient10−3
Learning rate1×10−6
Learning-rate scheduleconstant
Gradient clip1.0
Prompts per batch64
Rollouts per prompt8
Micro-batch per GPU1
Rollout temperature1.0
Rollout top-p1.0
Validation temperature0.7
Validation samples per instance1
Prompt budget30k tokens
Response budget170k tokens
Staleness threshold1
Partial-rollout recoveryon
Importance-sampling correctionoff
Training pool2,699 tasks
Epochs3
Figure 11: Reasoning share of the response over training on a fixed 120-task validation cohort. (a) Mean over tasks (solid) and character-weighted mean (dashed); (b) Median share at each task’s first (open) and last (filled) sampled epoch, grouped by number of rollouts solved.
Figure 11: Reasoning share of the response over training on a fixed 120-task validation cohort. (a) Mean over tasks (solid) and character-weighted mean (dashed); (b) Median share at each task’s first (open) and last (filled) sampled epoch, grouped by number of rollouts solved.

为什么重要

这让已经搭建好编程智能体系统的团队无需重写底层代码就能用强化学习提升性能,同时能及时发现原本会悄悄破坏训练效果的隐藏故障。随着智能体编程工具从研究演示走向生产级规模训练,这种可靠性和可观测性变得越来越重要。

Figure 12: Agent behaviors before and after training, computed over 420 trajectories at each end of the production OpenHands SDK run. The pass@k and passk rows are computed over prompt groups in the first and last third of the run.
Figure 12: Agent behaviors before and after training, computed over 420 trajectories at each end of the production OpenHands SDK run. The pass@k and passk rows are computed over prompt groups in the first and last third of the run.

本文术语

  • 强化学习(RL) · 通过对结果给予奖励来不断改进系统行为方式的训练方法
  • harness(运行框架) · 管理AI智能体工具调用、上下文和执行反馈的软件程序
  • 沙盒(sandbox) · 用于安全运行和测试代码的隔离环境
  • 混合专家模型(MoE) · 一种大模型结构,每次只激活其中部分子网络(专家)
  • SWE-bench Verified · 用于检验智能体解决真实软件仓库问题能力的基准测试

论文原文摘要(英文)

Reinforcement learning for coding agents increasingly relies on long-running agent harnesses to manage tool integration, repository contexts, and execution feedback. However, the native execution environments of these harnesses are inherently misaligned with policy-gradient training: environmental crashes and reward hacking corrupt outcome signals, while train-inference discrepancies decouple rollout behavior from policy updates. To address this, we present LEGO-RL, a framework that bridges native coding-agent harnesses with scalable policy-gradient optimization without modifying their internal control flow. LEGO-RL is built upon three pillars: (1) faithful optimization via in-process LLM proxying that captures raw generation streams for token-level alignment and robust trainer-side log-probability recomputation, even under harness-side compaction or re-serialization; (2) reliable execution via scalable sandbox orchestration featuring image caching and stage-wise defenses to mitigate reward hacking; and (3) observable training through an integrated plugin that automates validation and monitoring, paired with a Live UI for granular trajectory diagnostics. We evaluate LEGO-RL by training the sparse MoE model Qwen3.5-35B-A3B with GSPO across three native coding-agent harnesses. LEGO-RL improves Qwen3.5-35B-A3B across OpenHands SDK (64.0% to 70.4%), Claude Code (62.4% to 68.2%), and OpenCode (57.2% to 66.6%) on SWE-bench Verified, while maintaining a rollout-training probability correlation above 0.99.

作者 · Yiming Du

在 arXiv 阅读

最新论文

全部论文 →

METAL MEDIA 最新报道

图片来源: Yiming Du et al., arXiv:2608.17393, CC0 1.0