Everything K-culture — comebacks to K-beauty, straight to your inboxGet it in your inbox

METAL MEDIA

Evidence-RL: Towards Evidence-intensive Visual Reasoning

arXiv:2608.080212026-08-07

A training method that checks whether a correct answer actually depends on the right part of the image, not just whether it's correct

Vision-Language Models can give correct-looking answers from language priors or dataset shortcuts without actually looking at the relevant image region. This paper introduces CED, a counterfactual audit that measures whether a sampled answer's support drops more when its evidence region is masked out than when unrelated regions are masked, and folds this signal into GRPO reinforcement learning as Evidence-RL. Across nine public benchmarks and four backbones, Evidence-RL outperformed prior RL-based post-training methods.

METAL MEDIA explanatory visual

Structure of the Evidence-RL Training Pipeline

Evidence statusMeasured results reported

  1. Sample candidate answerThe policy model generates candidate answers to an image-question pair
  2. Evidence vs. non-evidence interventionVisual tokens inside a COCO-derived Evidence Region, and inside matched non-evidence regions, are each replaced with the mean of neighboring tokens
  3. Compute evidence marginThe drop in answer support from the evidence-region intervention is compared against the average drop from non-evidence regions to score how much the answer relies on real evidence
  4. Combine into GRPO rewardThe evidence margin is multiplied into the correctness score, so equally correct answers get different rewards based on evidence dependence
  5. InferenceThe trained model runs without this audit process at inference time, so no extra computation is added at deployment
An explanatory diagram made by METAL MEDIA, not a figure supplied by the paper's authors.

What they did

  1. Motivation: prior perception-aware RL methods (PAPO, VPPO) test coarse image sensitivity via global image corruption or attention proxies, but never check whether the specific answer causally depends on the local visual evidence that should support it.
  2. Method: for each sampled answer, CED replaces the visual tokens of an object-centric 'Evidence Region' (from COCO boxes) with the mean of neighboring tokens, measures the drop in support for that answer, and compares it against the same intervention applied to matched non-evidence regions to get an 'evidence margin'.
  3. This evidence margin is combined with answer correctness inside GRPO reward, so among equally correct rollouts, the one relying on real visual evidence gets a much higher reward — in one case study two rollouts both answered '4' correctly but the evidence-grounded one received a 7x larger reward than the prior-based one.
  4. Evaluated on nine benchmarks (CountBench, HallusionBench, MMMU, etc.) across four backbones (Qwen2.5-VL, Qwen3-VL, Qwen3.5-9B, plus reference comparisons with LLaVA and InternVL3.5), plus ablations comparing answer-only vs chain-of-thought scoring, intervention types (mean vs zero vs noise replacement), and robustness to imperfect proposal boxes.
Figure 1: Left: Task setup with a counting query over two views of the same scene. Upper right: A pretrained VLM gives the same prior-based answer for both views; Evidence-RL inspects the visual evidence and adjusts its answer when the scene changes. Lower right: Causal path decomposition: a candidate answer can be supported by the target visual evidence, by irrelevant context, or by language priors, which motivates CED to test whether the response depends specifically on the evidence path.
Figure 1: Left: Task setup with a counting query over two views of the same scene. Upper right: A pretrained VLM gives the same prior-based answer for both views; Evidence-RL inspects the visual evidence and adjusts its answer when the scene changes. Lower right: Causal path decomposition: a candidate answer can be supported by the target visual evidence, by irrelevant context, or by language priors, which motivates CED to test whether the response depends specifically on the evidence path.
Table 1: Signal-validation statistics on counting and presence tasks.
TaskNon-constant raw rewardZero varianceMean reward stdSame answer, different rewardMean negation
Counting99.5%24.0%0.129990.0%0.003
Presence21.1%79.0%0.04975.0%0.777
Figure 2: Overview of Evidence-RL during post-training. A policy samples candidate answers, and CED scores each answer by applying the same feature-space intervention to the proposed Evidence Region and to matched non-evidence Regions. The contrastive support drop yields an evidence margin, which is combined with correctness to form the GRPO reward. The counterfactual audit is used only during training; inference uses the trained VLM normally.
Figure 2: Overview of Evidence-RL during post-training. A policy samples candidate answers, and CED scores each answer by applying the same feature-space intervention to the proposed Evidence Region and to matched non-evidence Regions. The contrastive support drop yields an evidence margin, which is combined with correctness to form the GRPO reward. The counterfactual audit is used only during training; inference uses the trained VLM normally.
Table 2: Nine-benchmark evaluation using Qwen2.5-VL-7B as the base model for Ours. Backbone, baseline, and benchmark sources are cited in the setup paragraph. Best bold; second-best underlined. Each Δ row reports absolute gain over the corresponding base.
GroundingGeneral Reasoning
ModelCountBenchSpatialEvalHallusionVLMs AreBlindFREAKMathVistaMMBenchMMMUScienceQAAvg
VLM backbones
LLaVA-v1.6-7B55.6023.6051.0028.239.5022.8071.3029.0057.2038.69
InternVL3.5-8B86.9024.0069.6054.9212.8051.9083.4050.8089.9058.25
Qwen2.5-VL-3B70.7153.9465.3742.4611.1452.1080.4050.2974.9355.70
Qwen2.5-VL-7B81.8259.2168.6446.4412.1863.1084.7350.6383.9861.19
RL-based methods
VAPO-Thinker-7B86.9060.8071.2048.6213.1051.1081.8044.6082.3060.05
Δ+5.10+1.59+2.60+2.18+0.92−12.00−2.90−6.00−1.70−1.14
VPPO-7B85.9061.8068.2049.3512.7067.9084.9052.1088.4063.47
Δ+4.10+2.59−0.40+2.91+0.52+4.80+0.20+1.50+4.40+2.28
Perception-R1-7B84.9059.7066.7045.2212.2067.1081.7048.1082.3060.88
Δ+3.10+0.49−1.90−1.22+0.02+4.00−3.00−2.50−1.70−0.31
PAPO-G-H90.9064.8069.5047.9212.8069.4082.8050.5085.6063.80
Δ+9.10+5.59+0.90+1.48+0.62+6.30−1.90−0.10+1.60+2.61
VLM-R174.8056.4066.6044.1511.0060.3079.9048.6073.6057.26
Δ vs. Qwen2.5-VL-3B+4.10+2.50+1.20+1.69−0.14+8.20−0.50−1.70−1.30+1.56
SophiaVL-R182.8361.3866.8747.0924.1866.3087.4152.1887.6463.99
Δ+1.01+2.17−1.77+0.65+12.00+3.20+2.68+1.55+3.66+2.80
Ours88.8963.3470.0854.0226.4069.4087.5357.4787.0867.13
Δ+7.07+4.13+1.44+7.58+14.22+6.30+2.80+6.84+3.10+5.94
Figure 3: A case showing why correctness-only reward is insufficient. Two rollouts from the same GRPO group both answer “4” correctly, but the CED gate g⁡(m) creates a 7× reward gap between a prior-based rollout (g⁡(m)=0.18, R=0.11) and an evidence-grounded rollout (g⁡(m)=1.00, R=0.78), enabling GRPO to select the grounded trajectory.
Figure 3: A case showing why correctness-only reward is insufficient. Two rollouts from the same GRPO group both answer “4” correctly, but the CED gate g⁡(m) creates a 7× reward gap between a prior-based rollout (g⁡(m)=0.18, R=0.11) and an evidence-grounded rollout (g⁡(m)=1.00, R=0.78), enabling GRPO to select the grounded trajectory.
Table 3: Controlled comparison on Qwen3.5-9B with matched data, router, and compute. Δ is the mean gain over the frozen base; ΔCED is Answer-CED minus correctness-only.
GroundingGeneral Reasoning
MethodCountBenchSpatialEvalHallusionVLMs AreBlindFREAKMathVistaMMBenchMMMUScienceQAAvgΔ
λ=0 (corr.-only)89.9040.4557.9372.9018.7373.4079.4436.0082.2461.22−1.24
VPPO (retrained)86.8738.3258.3772.9719.0772.5079.6237.3381.0960.68−1.78
PAPO (retrained)†86.878.5958.4669.9619.4014.0044.8733.2234.2441.07−21.39
Answer-CED93.9054.3578.1073.1023.1283.7588.5875.6793.6173.80+11.34
ΔCED+4.00+13.90+20.17+0.20+4.39+10.35+9.14+39.67+11.37+12.58
Figure 4: Training dynamics for Answer-CED and CoT-CED on Qwen3.5-9B. CoT-CED produces larger training-side rewards, but the downstream comparison in Table 5 favors Answer-CED on average; we therefore use Answer-CED as the default variant.
Figure 4: Training dynamics for Answer-CED and CoT-CED on Qwen3.5-9B. CoT-CED produces larger training-side rewards, but the downstream comparison in Table 5 favors Answer-CED on average; we therefore use Answer-CED as the default variant.
Table 4: Cross-backbone validation of Ours. Each block reports the frozen base, Ours, and absolute Δ against that block’s base. The Qwen3.5-9B block additionally reports 3-seed mean±std, with its mean Δ in the last column.
GroundingGeneral Reasoning
ModelCountBenchSpatialEvalHallusionVLMs AreBlindFREAKMathVistaMMBenchMMMUScienceQAAvg. Δ
Backbone: Qwen2.5-VL-3B
Qwen2.5-VL-3B70.7153.9465.3742.4611.1452.1080.4050.2974.93
Ours70.7155.5667.4049.5125.5160.4082.5550.5780.43
Δ+0.00+1.62+2.03+7.05+14.37+8.30+2.15+0.28+5.50+4.59
Backbone: Qwen2.5-VL-7B
Qwen2.5-VL-7B81.8259.2168.6446.4412.1863.1084.7350.6383.98
Ours88.8963.3470.0854.0226.4069.4087.5357.4787.08
Δ+7.07+4.13+1.44+7.58+14.22+6.30+2.80+6.84+3.10+5.94
Backbone: Qwen3-VL-8B-Instruct
Qwen3-VL-8B94.9065.2672.8667.0128.8567.5489.1557.8291.99
Ours95.9267.7574.3068.8831.1868.0489.5858.8592.82
Δ+1.02+2.49+1.44+1.87+2.33+0.50+0.43+1.03+0.83+1.33
Backbone: Qwen3.5-9B
Qwen3.5-9B80.8052.7347.2046.449.9580.0088.0366.2190.78
Ours93.9054.3578.1073.1023.1283.7588.5875.6793.61
Δ+13.10+1.62+30.90+26.66+13.17+3.75+0.55+9.46+2.83+11.34
Ours (3 seeds)93.60 ±0.4856.04 ±0.2778.35 ±0.2173.37 ±0.1723.11 ±0.4084.40 ±0.5089.87 ±0.0475.39 ±0.2294.46 ±0.13+11.83
Figure 5: CED probe robustness under proposal perturbation. (a) Replacing the COCO proposal with a random object box as Ωev collapses the evidence margin to near zero across all seven task types, confirming that the signal tracks proposal relevance rather than masking magnitude. (b) IoU-graded degradation on the count-exclusion diagnostic: m¯ decreases monotonically as spatial overlap with the original proposal drops, approaching the random-region floor at lowest IoU.
Figure 5: CED probe robustness under proposal perturbation. (a) Replacing the COCO proposal with a random object box as Ωev collapses the evidence margin to near zero across all seven task types, confirming that the signal tracks proposal relevance rather than masking magnitude. (b) IoU-graded degradation on the count-exclusion diagnostic: m¯ decreases monotonically as spatial overlap with the original proposal drops, approaching the random-region floor at lowest IoU.
Table 5: Answer vs. CoT variants on Qwen3.5-9B; Avg. Δ is the mean gain over the base.
GroundingGeneral Reasoning
VariantCountBenchSpatialEvalHallusionVLMs AreBlindFREAKMathVistaMMBenchMMMUScienceQAAvg. Δ
Answer93.9054.3578.1073.1023.1283.7588.5875.6793.61+11.34
CoT92.9053.8078.1069.0930.3581.8888.5368.2394.66+10.60
Figure A.1: CED probe robustness under proposal perturbation on the count-exclusion diagnostic. (a) Mean evidence margin m¯ across IoU bins, with the random-region baseline as a grey dashed reference. (b,c) m¯ under scale and translation perturbations, with the original ±1 SE band in red. Error bars are ±1 SE.
Figure A.1: CED probe robustness under proposal perturbation on the count-exclusion diagnostic. (a) Mean evidence margin m¯ across IoU bins, with the random-region baseline as a grey dashed reference. (b,c) m¯ under scale and translation perturbations, with the original ±1 SE band in red. Error bars are ±1 SE.
Table A.1: Complete hyperparameter listing. “Paper notation” gives the corresponding symbol in the main text when applicable.
CategoryParameterValuePaper notation
RL & optimization
Learning ratelr1×10−5
KL penalty coefficientkl_coeff0.01
GRPO group sizegroup_size32
Training stepsn_steps2,000
Trainable layersn_trainable_layers4 (last)
Precisiondtypebfloat16
Generation & sampling
Temperaturetemperature1.0
Top-ptop_p0.95
Max new tokens (evidence)max_new_tokens32
Max response tokens (full)max_response_tokens64
Prompt formatshort_evidence_v1
Reward function (Equation 4)
Response reward weightalpha_resp0.70see note†
Answer reward weightalpha_ans0.30see note†
Response logprob temp.tau_resp0.20
Answer logprob temp.tau_ans1.00
Gate temperaturetau_resp (shared)0.20τg
Gate baseline at m=0(derived)0.50g⁡(0)
Evidence tie-breakerevidence_eps0.10εtie
Additive evidence weightevidence_eps (shared)0.10λ
Reward clipping[min, max][−1.25, 1.00]
Format penalty−0.75
Data & evaluation
Rerank candidatesrerank_num_candidates6
non-evidence Regions per sampleK3K
Figure A.2: Qualitative examples of same-answer-different-reward under the CED method. All rollouts share the same answer, while CED produces reward spans of 0.50–0.67 according to visual evidence quality.
Figure A.2: Qualitative examples of same-answer-different-reward under the CED method. All rollouts share the same answer, while CED produces reward spans of 0.50–0.67 according to visual evidence quality.
Table A.2: Methodological comparison with perception-aware VLM RL methods. The Causal audit column indicates whether the method tests causal evidence dependence for a candidate answer.
MethodSignal sourceGranularityCausal auditStage
PAPO [36]Global mask KLGlobalNoTraining
VPPO [10]Attention weightsToken-levelNoTraining
Vision-SR1†Text self-verificationDescription-levelNoTraining
PaLMR‡LLM-as-JudgeTrajectory-levelNoTraining
SophiaVL-R1 [6]Thinking-reward RMTrajectory-levelNoTraining
PEARL§Perception checklistSample-levelNoTraining
VAPO¶Trajectory anchoringToken-levelNoTraining
VLM-R1*Self-verificationResponse-levelNoTraining
Perception-R1††Proxy localizationRegion-levelaNoTraining
CED (Ours)Counterfactual interventionRegion-levelYesTraining
Figure A.3: Within-group zero-variance rate across task families. Evidence-intensive tasks show a large reduction from correctness-only reward to our reward; binary-dominated groups remain near high zero-variance because of their two-action structure.
Figure A.3: Within-group zero-variance rate across task families. Evidence-intensive tasks show a large reduction from correctness-only reward to our reward; binary-dominated groups remain near high zero-variance because of their two-action structure.
Table A.3: Cross-model evidence sensitivity validation (N=50 per task per model).
ModelTaskRelevant > Random rateMean margin
InternVL3.5-8BCounting0.680.0070
Presence0.620.0737
LLaVA-v1.6-7BCounting0.780.0074
Presence0.600.1625
Figure A.4: Failure cases on presence tasks: identical negation answers with identical rewards yield zero within-group variance.
Figure A.4: Failure cases on presence tasks: identical negation answers with identical rewards yield zero within-group variance.
Table A.4: Blindfold test results (N=500 per task family).
ConditionTaskRel. > Rand. rateMean Delta
NormalCounting0.634+1.928
NormalPresence0.550+0.780
BlindfoldCounting0.404−0.114
BlindfoldPresence0.448+0.070
Table A.5: Perturbation-family ablation within the local intervention family, evaluated by logits-JS AUC.
Key modeMean repl.Zero repl.Gaussian-noise repl.
prompt_last0.67800.66020.6299
answer_first0.65920.62150.6273
Average0.66860.64090.6286
Table A.6: Gaussian-noise profile under prompt_last. Stability drops under logits_only.
ConfigFixed metricBest metricBest paired AUC
logits240.62990.62990.7108
logits_only0.57530.59200.6407
Table A.7: Evidence weight (λ) sensitivity. λ=0.00 is correctness-only.
λNon-const. rateReward stdZero-var. rate
0.00100%0.05280%
0.05100%0.05080%
0.1096%0.05974%
0.15100%0.04154%
0.20100%0.04910%
Table A.8: Gate temperature (τg) sensitivity.
τgNon-const. rateReward stdZero-var. rate
0.10100%0.05684%
0.2096%0.05974%
0.3096%0.04694%
0.50100%0.03780%
Table A.9: Proposal-perturbation results on count-exclusion (n=41). Higher values indicate stronger evidence dependence; the random row gives the floor without spatial targeting.
AxisConditionΔRel. rate
IoU stratification[0.7, 1.0]0.5012.0320.732
[0.5, 0.7)0.3791.2820.683
[0.3, 0.5)0.3411.3860.683
[0.0, 0.3)0.2190.5550.561
Scale0.50×0.3771.1130.732
0.75×0.4811.7500.732
1.25×0.4332.2300.683
1.50×0.2952.0030.634
2.00×0.3821.8800.707
Shift5%0.3041.6020.634
10%0.4281.4570.707
20%0.5321.8990.732
original proposal (reference)0.3861.8520.683
random region (floor)0.1470.2770.537
Table A.10: Unified-condition baseline comparison on the offline rerank dataset (N=1,353).
MethodReward meanSignal strengthSignal-corr
Vanilla GRPO−0.149N/AN/A
PAPO-Lite−0.156KL =0.0240.034
CED (Ours)+0.448Margin =−0.007Selective
Table A.11: Unified-condition method comparison on the counting shared-candidate slice: 94 samples with at least two distinct candidates.
MethodAcc.Strong inf.rnc↑flip ↑
LogProb94.775.0
Correctness-only97.9100.012.84.3
CED (Ours)97.9100.0100.076.6
Perception-R1-Lite97.9100.013.85.3
VLM-R1-Lite∗94.775.00.00.0
VPPO-Lite∗94.775.00.00.0
PAPO-Lite∗94.775.00.00.0
∗ No reward-based reranking recovered; falls back to log probability.
Table A.12: CoT-CED vs. Answer-CED: training-end summary over 2,000 steps.
MetricCoT-CEDAnswer-CED
Mean training reward+0.259−0.009
Final training correctness43.4%34.2%
Table A.13: End-of-training CoT behavior on held-out prompts. CoT-CED collapses toward an object-cue answer format, while Answer-CED preserves natural multi-sentence reasoning.
MetricBenchmarkCoT-CEDAnswer-CED
Mean CoT length (tokens)Counting3.649.3
Median CoT length (tokens)Counting347
p90 CoT length (tokens)Counting678
Mean CoT length (tokens)HallusionBench3.882.1
Median CoT length (tokens)HallusionBench361
p90 CoT length (tokens)HallusionBench6173
Image-referential opener rateCounting1.0%79.8%
Image-referential opener rateHallusionBench5.2%64.2%
Unique-token ratioCounting99.1%75.8%
Unique-token ratioHallusionBench96.8%68.1%
Table A.14: Representative side-by-side generations. Both variants produce the correct final answer, but CoT-CED reduces the chain to an object cue, while Answer-CED keeps an explicit visual check.
PromptCoT-CEDAnswer-CED
How many beds are there in the image?The beds / 2Got it, let’s look at the image. There’s a loft bed with a lower bed. The upper bunk is one, the lower bed is another. So that’s two beds.
How many weights are there in the image?The weights / TwoGot it, let’s look at the image. The man is holding two dumbbells, one in each hand. There’s one in the left hand and one in the right hand. So that’s two weights.

Findings

  • Table 3: under matched conditions on Qwen3.5-9B (same data, router, 2,000-step schedule), Answer-CED improved the average score by +12.58 over correctness-only reward and by +13.12 over a retrained VPPO baseline.
  • Table 2: using Qwen2.5-VL-7B as base, Evidence-RL achieved the highest average score and largest mean gain among compared RL-based post-training methods across nine benchmarks, and was the only method with non-negative gains on all nine.
  • Table 4: across four backbones from two model families (Qwen2.5-VL-3B/7B, Qwen3-VL-8B, Qwen3.5-9B), the method produced positive mean improvement on every backbone and non-negative gains on all 36 benchmark-backbone combinations.
  • Figure 5: replacing the COCO-derived proposal with a random object box collapsed the evidence margin from 0.268 to 0.015, and the relevant-beats-random rate dropped from 0.636 to 0.487 (chance level), confirming the signal tracks proposal relevance rather than masking magnitude alone.
  • On eight text-only benchmarks (ARC, OpenBookQA, CommonsenseQA, etc.), the Qwen3.5-9B Answer-CED checkpoint showed a mean accuracy change of only -0.27 percentage points, with no single benchmark dropping more than 1.5 points.

Where it can be used

  • Reward design for RL post-training of VLMs on VQA, counting, or spatial reasoning tasks where hallucination or shortcut reasoning is a concern
  • Diagnosing or filtering training data to separate 'guessed correctly' answers from 'evidence-grounded' answers when accuracy alone can't distinguish them
  • Settings needing a low-cost visual grounding audit signal built from weak object-detection boxes (e.g., COCO-style annotations) without question-specific evidence labels

Limits and open work

  • Experiments are limited to object-centric region interventions using COCO-style boxes; extension to attribute-level, relational, or text-span evidence is proposed conceptually but not separately validated.
  • Binary yes/no presence-type questions show inherently limited within-group reward variance due to their two-action structure, a reported constraint of the approach.
  • The counterfactual audit is used only during training; at inference time the trained model runs normally without this check, so whether grounding persists in deployment is a separate question not directly measured here.
  • Scale and translation perturbations to the proposal region showed no systematic degradation only within the tested range, so robustness beyond that range is unverified.
  • On blank-placeholder-image ScienceQA items, the CED model correctly refuses to answer ('cannot determine from image') but is scored incorrect against a benchmark that rewards prior-based guessing, showing a scoring mismatch on that subset.

Why it matters

Accuracy-only evaluation can't tell apart a model that guessed correctly from one that actually looked at the image, and this method gives a practical way to reward the difference during training. It matters for anyone building VQA, counting, or spatial-reasoning systems where hallucination or shortcut reasoning from language priors is a known failure mode.

Terms in this paper

  • VLM (Vision-Language Model) · a model that takes both an image and text as input to produce answers
  • counterfactual intervention · artificially removing information from a specific region and observing how the model's output changes
  • GRPO · a reinforcement learning method that samples multiple answers per question and rewards them relative to each other within the group
  • evidence margin · the difference between how much support drops when the evidence region is masked versus when other regions are masked
  • Answer-CED vs CoT-CED · two variants of the method — one scores only the final answer span, the other scores the whole reasoning trace

Original abstract (English)

Vision-Language Models (VLMs) should answer from concrete image evidence rather than language priors, dataset shortcuts, or irrelevant visual context. Existing perception-aware post-training methods encourage image use through global perturbations or attention proxies, but they do not test whether a sampled answer causally depends on the local evidence that supports it. We propose Counterfactual Evidence Disentanglement (CED), a training-time evidence audit for VLM grounding. For each response, CED neutralizes an object-centric Evidence Region and compares the resulting support drop against matched non-evidence Regions. We combine this signal with answer correctness inside GRPO, rewarding correct answers that rely on the evidence path rather than shortcut or nuisance paths. CED uses weak object-level proposals, requires no question-specific evidence annotations, and adds no inference-time overhead. Across nine public benchmarks and four backbones, CED outperforms prior RL-based post-training methods, with targeted analyses verifying its object-centric signal.

Authors · Haojie Huang

Read on arXiv

Latest papers

All papers →

Latest from METAL MEDIA

Figures: Haojie Huang et al., arXiv:2608.08021, arxiv-nonexclusive