Everything K-culture — comebacks to K-beauty, straight to your inboxGet it in your inbox›
Evidence-RL: Towards Evidence-intensive Visual Reasoning
arXiv:2608.080212026-08-07
A training method that checks whether a correct answer actually depends on the right part of the image, not just whether it's correct
Vision-Language Models can give correct-looking answers from language priors or dataset shortcuts without actually looking at the relevant image region. This paper introduces CED, a counterfactual audit that measures whether a sampled answer's support drops more when its evidence region is masked out than when unrelated regions are masked, and folds this signal into GRPO reinforcement learning as Evidence-RL. Across nine public benchmarks and four backbones, Evidence-RL outperformed prior RL-based post-training methods.
METAL MEDIA explanatory visual
Structure of the Evidence-RL Training Pipeline
Evidence statusMeasured results reported
Sample candidate answerThe policy model generates candidate answers to an image-question pair
Evidence vs. non-evidence interventionVisual tokens inside a COCO-derived Evidence Region, and inside matched non-evidence regions, are each replaced with the mean of neighboring tokens
Compute evidence marginThe drop in answer support from the evidence-region intervention is compared against the average drop from non-evidence regions to score how much the answer relies on real evidence
Combine into GRPO rewardThe evidence margin is multiplied into the correctness score, so equally correct answers get different rewards based on evidence dependence
InferenceThe trained model runs without this audit process at inference time, so no extra computation is added at deployment
An explanatory diagram made by METAL MEDIA, not a figure supplied by the paper's authors.
What they did
Motivation: prior perception-aware RL methods (PAPO, VPPO) test coarse image sensitivity via global image corruption or attention proxies, but never check whether the specific answer causally depends on the local visual evidence that should support it.
Method: for each sampled answer, CED replaces the visual tokens of an object-centric 'Evidence Region' (from COCO boxes) with the mean of neighboring tokens, measures the drop in support for that answer, and compares it against the same intervention applied to matched non-evidence regions to get an 'evidence margin'.
This evidence margin is combined with answer correctness inside GRPO reward, so among equally correct rollouts, the one relying on real visual evidence gets a much higher reward — in one case study two rollouts both answered '4' correctly but the evidence-grounded one received a 7x larger reward than the prior-based one.
Evaluated on nine benchmarks (CountBench, HallusionBench, MMMU, etc.) across four backbones (Qwen2.5-VL, Qwen3-VL, Qwen3.5-9B, plus reference comparisons with LLaVA and InternVL3.5), plus ablations comparing answer-only vs chain-of-thought scoring, intervention types (mean vs zero vs noise replacement), and robustness to imperfect proposal boxes.
Figure 1: Left: Task setup with a counting query over two views of the same scene. Upper right: A pretrained VLM gives the same prior-based answer for both views; Evidence-RL inspects the visual evidence and adjusts its answer when the scene changes. Lower right: Causal path decomposition: a candidate answer can be supported by the target visual evidence, by irrelevant context, or by language priors, which motivates CED to test whether the response depends specifically on the evidence path.
Table 1: Signal-validation statistics on counting and presence tasks.
Task
Non-constant raw reward
Zero variance
Mean reward std
Same answer, different reward
Mean negation
Counting
99.5%
24.0%
0.1299
90.0%
0.003
Presence
21.1%
79.0%
0.0497
5.0%
0.777
Figure 2: Overview of Evidence-RL during post-training. A policy samples candidate answers, and CED scores each answer by applying the same feature-space intervention to the proposed Evidence Region and to matched non-evidence Regions. The contrastive support drop yields an evidence margin, which is combined with correctness to form the GRPO reward. The counterfactual audit is used only during training; inference uses the trained VLM normally.
Table 2: Nine-benchmark evaluation using Qwen2.5-VL-7B as the base model for Ours. Backbone, baseline, and benchmark sources are cited in the setup paragraph. Best bold; second-best underlined. Each Δ row reports absolute gain over the corresponding base.
Grounding
General Reasoning
Model
CountBench
SpatialEval
Hallusion
VLMs AreBlind
FREAK
MathVista
MMBench
MMMU
ScienceQA
Avg
VLM backbones
LLaVA-v1.6-7B
55.60
23.60
51.00
28.23
9.50
22.80
71.30
29.00
57.20
38.69
InternVL3.5-8B
86.90
24.00
69.60
54.92
12.80
51.90
83.40
50.80
89.90
58.25
Qwen2.5-VL-3B
70.71
53.94
65.37
42.46
11.14
52.10
80.40
50.29
74.93
55.70
Qwen2.5-VL-7B
81.82
59.21
68.64
46.44
12.18
63.10
84.73
50.63
83.98
61.19
RL-based methods
VAPO-Thinker-7B
86.90
60.80
71.20
48.62
13.10
51.10
81.80
44.60
82.30
60.05
Δ
+5.10
+1.59
+2.60
+2.18
+0.92
−12.00
−2.90
−6.00
−1.70
−1.14
VPPO-7B
85.90
61.80
68.20
49.35
12.70
67.90
84.90
52.10
88.40
63.47
Δ
+4.10
+2.59
−0.40
+2.91
+0.52
+4.80
+0.20
+1.50
+4.40
+2.28
Perception-R1-7B
84.90
59.70
66.70
45.22
12.20
67.10
81.70
48.10
82.30
60.88
Δ
+3.10
+0.49
−1.90
−1.22
+0.02
+4.00
−3.00
−2.50
−1.70
−0.31
PAPO-G-H
90.90
64.80
69.50
47.92
12.80
69.40
82.80
50.50
85.60
63.80
Δ
+9.10
+5.59
+0.90
+1.48
+0.62
+6.30
−1.90
−0.10
+1.60
+2.61
VLM-R1
74.80
56.40
66.60
44.15
11.00
60.30
79.90
48.60
73.60
57.26
Δ vs. Qwen2.5-VL-3B
+4.10
+2.50
+1.20
+1.69
−0.14
+8.20
−0.50
−1.70
−1.30
+1.56
SophiaVL-R1
82.83
61.38
66.87
47.09
24.18
66.30
87.41
52.18
87.64
63.99
Δ
+1.01
+2.17
−1.77
+0.65
+12.00
+3.20
+2.68
+1.55
+3.66
+2.80
Ours
88.89
63.34
70.08
54.02
26.40
69.40
87.53
57.47
87.08
67.13
Δ
+7.07
+4.13
+1.44
+7.58
+14.22
+6.30
+2.80
+6.84
+3.10
+5.94
Figure 3: A case showing why correctness-only reward is insufficient. Two rollouts from the same GRPO group both answer “4” correctly, but the CED gate g(m) creates a 7× reward gap between a prior-based rollout (g(m)=0.18, R=0.11) and an evidence-grounded rollout (g(m)=1.00, R=0.78), enabling GRPO to select the grounded trajectory.
Table 3: Controlled comparison on Qwen3.5-9B with matched data, router, and compute. Δ is the mean gain over the frozen base; ΔCED is Answer-CED minus correctness-only.
Grounding
General Reasoning
Method
CountBench
SpatialEval
Hallusion
VLMs AreBlind
FREAK
MathVista
MMBench
MMMU
ScienceQA
Avg
Δ
λ=0 (corr.-only)
89.90
40.45
57.93
72.90
18.73
73.40
79.44
36.00
82.24
61.22
−1.24
VPPO (retrained)
86.87
38.32
58.37
72.97
19.07
72.50
79.62
37.33
81.09
60.68
−1.78
PAPO (retrained)†
86.87
8.59
58.46
69.96
19.40
14.00
44.87
33.22
34.24
41.07
−21.39
Answer-CED
93.90
54.35
78.10
73.10
23.12
83.75
88.58
75.67
93.61
73.80
+11.34
ΔCED
+4.00
+13.90
+20.17
+0.20
+4.39
+10.35
+9.14
+39.67
+11.37
+12.58
–
Figure 4: Training dynamics for Answer-CED and CoT-CED on Qwen3.5-9B. CoT-CED produces larger training-side rewards, but the downstream comparison in Table 5 favors Answer-CED on average; we therefore use Answer-CED as the default variant.
Table 4: Cross-backbone validation of Ours. Each block reports the frozen base, Ours, and absolute Δ against that block’s base. The Qwen3.5-9B block additionally reports 3-seed mean±std, with its mean Δ in the last column.
Grounding
General Reasoning
Model
CountBench
SpatialEval
Hallusion
VLMs AreBlind
FREAK
MathVista
MMBench
MMMU
ScienceQA
Avg. Δ
Backbone: Qwen2.5-VL-3B
Qwen2.5-VL-3B
70.71
53.94
65.37
42.46
11.14
52.10
80.40
50.29
74.93
–
Ours
70.71
55.56
67.40
49.51
25.51
60.40
82.55
50.57
80.43
–
Δ
+0.00
+1.62
+2.03
+7.05
+14.37
+8.30
+2.15
+0.28
+5.50
+4.59
Backbone: Qwen2.5-VL-7B
Qwen2.5-VL-7B
81.82
59.21
68.64
46.44
12.18
63.10
84.73
50.63
83.98
–
Ours
88.89
63.34
70.08
54.02
26.40
69.40
87.53
57.47
87.08
–
Δ
+7.07
+4.13
+1.44
+7.58
+14.22
+6.30
+2.80
+6.84
+3.10
+5.94
Backbone: Qwen3-VL-8B-Instruct
Qwen3-VL-8B
94.90
65.26
72.86
67.01
28.85
67.54
89.15
57.82
91.99
–
Ours
95.92
67.75
74.30
68.88
31.18
68.04
89.58
58.85
92.82
–
Δ
+1.02
+2.49
+1.44
+1.87
+2.33
+0.50
+0.43
+1.03
+0.83
+1.33
Backbone: Qwen3.5-9B
Qwen3.5-9B
80.80
52.73
47.20
46.44
9.95
80.00
88.03
66.21
90.78
–
Ours
93.90
54.35
78.10
73.10
23.12
83.75
88.58
75.67
93.61
–
Δ
+13.10
+1.62
+30.90
+26.66
+13.17
+3.75
+0.55
+9.46
+2.83
+11.34
Ours (3 seeds)
93.60 ±0.48
56.04 ±0.27
78.35 ±0.21
73.37 ±0.17
23.11 ±0.40
84.40 ±0.50
89.87 ±0.04
75.39 ±0.22
94.46 ±0.13
+11.83
Figure 5: CED probe robustness under proposal perturbation. (a) Replacing the COCO proposal with a random object box as Ωev collapses the evidence margin to near zero across all seven task types, confirming that the signal tracks proposal relevance rather than masking magnitude. (b) IoU-graded degradation on the count-exclusion diagnostic: m¯ decreases monotonically as spatial overlap with the original proposal drops, approaching the random-region floor at lowest IoU.
Table 5: Answer vs. CoT variants on Qwen3.5-9B; Avg. Δ is the mean gain over the base.
Grounding
General Reasoning
Variant
CountBench
SpatialEval
Hallusion
VLMs AreBlind
FREAK
MathVista
MMBench
MMMU
ScienceQA
Avg. Δ
Answer
93.90
54.35
78.10
73.10
23.12
83.75
88.58
75.67
93.61
+11.34
CoT
92.90
53.80
78.10
69.09
30.35
81.88
88.53
68.23
94.66
+10.60
Figure A.1: CED probe robustness under proposal perturbation on the count-exclusion diagnostic. (a) Mean evidence margin m¯ across IoU bins, with the random-region baseline as a grey dashed reference. (b,c) m¯ under scale and translation perturbations, with the original ±1 SE band in red. Error bars are ±1 SE.
Table A.1: Complete hyperparameter listing. “Paper notation” gives the corresponding symbol in the main text when applicable.
Category
Parameter
Value
Paper notation
RL & optimization
Learning rate
lr
1×10−5
—
KL penalty coefficient
kl_coeff
0.01
—
GRPO group size
group_size
32
—
Training steps
n_steps
2,000
—
Trainable layers
n_trainable_layers
4 (last)
—
Precision
dtype
bfloat16
—
Generation & sampling
Temperature
temperature
1.0
—
Top-p
top_p
0.95
—
Max new tokens (evidence)
max_new_tokens
32
—
Max response tokens (full)
max_response_tokens
64
—
Prompt format
short_evidence_v1
—
—
Reward function (Equation 4)
Response reward weight
alpha_resp
0.70
see note†
Answer reward weight
alpha_ans
0.30
see note†
Response logprob temp.
tau_resp
0.20
—
Answer logprob temp.
tau_ans
1.00
—
Gate temperature
tau_resp (shared)
0.20
τg
Gate baseline at m=0
(derived)
0.50
g(0)
Evidence tie-breaker
evidence_eps
0.10
εtie
Additive evidence weight
evidence_eps (shared)
0.10
λ
Reward clipping
[min, max]
[−1.25, 1.00]
—
Format penalty
—
−0.75
—
Data & evaluation
Rerank candidates
rerank_num_candidates
6
—
non-evidence Regions per sample
K
3
K
Figure A.2: Qualitative examples of same-answer-different-reward under the CED method. All rollouts share the same answer, while CED produces reward spans of 0.50–0.67 according to visual evidence quality.
Table A.2: Methodological comparison with perception-aware VLM RL methods. The Causal audit column indicates whether the method tests causal evidence dependence for a candidate answer.
Method
Signal source
Granularity
Causal audit
Stage
PAPO [36]
Global mask KL
Global
No
Training
VPPO [10]
Attention weights
Token-level
No
Training
Vision-SR1†
Text self-verification
Description-level
No
Training
PaLMR‡
LLM-as-Judge
Trajectory-level
No
Training
SophiaVL-R1 [6]
Thinking-reward RM
Trajectory-level
No
Training
PEARL§
Perception checklist
Sample-level
No
Training
VAPO¶
Trajectory anchoring
Token-level
No
Training
VLM-R1*
Self-verification
Response-level
No
Training
Perception-R1††
Proxy localization
Region-levela
No
Training
CED (Ours)
Counterfactual intervention
Region-level
Yes
Training
Figure A.3: Within-group zero-variance rate across task families. Evidence-intensive tasks show a large reduction from correctness-only reward to our reward; binary-dominated groups remain near high zero-variance because of their two-action structure.
Table A.3: Cross-model evidence sensitivity validation (N=50 per task per model).
Model
Task
Relevant > Random rate
Mean margin
InternVL3.5-8B
Counting
0.68
0.0070
Presence
0.62
0.0737
LLaVA-v1.6-7B
Counting
0.78
0.0074
Presence
0.60
0.1625
Figure A.4: Failure cases on presence tasks: identical negation answers with identical rewards yield zero within-group variance.
Table A.4: Blindfold test results (N=500 per task family).
Condition
Task
Rel. > Rand. rate
Mean Delta
Normal
Counting
0.634
+1.928
Normal
Presence
0.550
+0.780
Blindfold
Counting
0.404
−0.114
Blindfold
Presence
0.448
+0.070
Table A.5: Perturbation-family ablation within the local intervention family, evaluated by logits-JS AUC.
Key mode
Mean repl.
Zero repl.
Gaussian-noise repl.
prompt_last
0.6780
0.6602
0.6299
answer_first
0.6592
0.6215
0.6273
Average
0.6686
0.6409
0.6286
Table A.6: Gaussian-noise profile under prompt_last. Stability drops under logits_only.
Config
Fixed metric
Best metric
Best paired AUC
logits24
0.6299
0.6299
0.7108
logits_only
0.5753
0.5920
0.6407
Table A.7: Evidence weight (λ) sensitivity. λ=0.00 is correctness-only.
λ
Non-const. rate
Reward std
Zero-var. rate
0.00
100%
0.0528
0%
0.05
100%
0.0508
0%
0.10
96%
0.0597
4%
0.15
100%
0.0415
4%
0.20
100%
0.0491
0%
Table A.8: Gate temperature (τg) sensitivity.
τg
Non-const. rate
Reward std
Zero-var. rate
0.10
100%
0.0568
4%
0.20
96%
0.0597
4%
0.30
96%
0.0469
4%
0.50
100%
0.0378
0%
Table A.9: Proposal-perturbation results on count-exclusion (n=41). Higher values indicate stronger evidence dependence; the random row gives the floor without spatial targeting.
Axis
Condition
m¯
Δ
Rel. rate
IoU stratification
[0.7, 1.0]
0.501
2.032
0.732
[0.5, 0.7)
0.379
1.282
0.683
[0.3, 0.5)
0.341
1.386
0.683
[0.0, 0.3)
0.219
0.555
0.561
Scale
0.50×
0.377
1.113
0.732
0.75×
0.481
1.750
0.732
1.25×
0.433
2.230
0.683
1.50×
0.295
2.003
0.634
2.00×
0.382
1.880
0.707
Shift
5%
0.304
1.602
0.634
10%
0.428
1.457
0.707
20%
0.532
1.899
0.732
original proposal (reference)
0.386
1.852
0.683
random region (floor)
0.147
0.277
0.537
Table A.10: Unified-condition baseline comparison on the offline rerank dataset (N=1,353).
Method
Reward mean
Signal strength
Signal-corr
Vanilla GRPO
−0.149
N/A
N/A
PAPO-Lite
−0.156
KL =0.024
0.034
CED (Ours)
+0.448
Margin =−0.007
Selective
Table A.11: Unified-condition method comparison on the counting shared-candidate slice: 94 samples with at least two distinct candidates.
Method
Acc.
Strong inf.
rnc↑
flip ↑
LogProb
94.7
75.0
–
–
Correctness-only
97.9
100.0
12.8
4.3
CED (Ours)
97.9
100.0
100.0
76.6
Perception-R1-Lite
97.9
100.0
13.8
5.3
VLM-R1-Lite∗
94.7
75.0
0.0
0.0
VPPO-Lite∗
94.7
75.0
0.0
0.0
PAPO-Lite∗
94.7
75.0
0.0
0.0
∗ No reward-based reranking recovered; falls back to log probability.
Table A.12: CoT-CED vs. Answer-CED: training-end summary over 2,000 steps.
Metric
CoT-CED
Answer-CED
Mean training reward
+0.259
−0.009
Final training correctness
43.4%
34.2%
Table A.13: End-of-training CoT behavior on held-out prompts. CoT-CED collapses toward an object-cue answer format, while Answer-CED preserves natural multi-sentence reasoning.
Metric
Benchmark
CoT-CED
Answer-CED
Mean CoT length (tokens)
Counting
3.6
49.3
Median CoT length (tokens)
Counting
3
47
p90 CoT length (tokens)
Counting
6
78
Mean CoT length (tokens)
HallusionBench
3.8
82.1
Median CoT length (tokens)
HallusionBench
3
61
p90 CoT length (tokens)
HallusionBench
6
173
Image-referential opener rate
Counting
1.0%
79.8%
Image-referential opener rate
HallusionBench
5.2%
64.2%
Unique-token ratio
Counting
99.1%
75.8%
Unique-token ratio
HallusionBench
96.8%
68.1%
Table A.14: Representative side-by-side generations. Both variants produce the correct final answer, but CoT-CED reduces the chain to an object cue, while Answer-CED keeps an explicit visual check.
Prompt
CoT-CED
Answer-CED
How many beds are there in the image?
The beds / 2
Got it, let’s look at the image. There’s a loft bed with a lower bed. The upper bunk is one, the lower bed is another. So that’s two beds.
How many weights are there in the image?
The weights / Two
Got it, let’s look at the image. The man is holding two dumbbells, one in each hand. There’s one in the left hand and one in the right hand. So that’s two weights.
Findings
Table 3: under matched conditions on Qwen3.5-9B (same data, router, 2,000-step schedule), Answer-CED improved the average score by +12.58 over correctness-only reward and by +13.12 over a retrained VPPO baseline.
Table 2: using Qwen2.5-VL-7B as base, Evidence-RL achieved the highest average score and largest mean gain among compared RL-based post-training methods across nine benchmarks, and was the only method with non-negative gains on all nine.
Table 4: across four backbones from two model families (Qwen2.5-VL-3B/7B, Qwen3-VL-8B, Qwen3.5-9B), the method produced positive mean improvement on every backbone and non-negative gains on all 36 benchmark-backbone combinations.
Figure 5: replacing the COCO-derived proposal with a random object box collapsed the evidence margin from 0.268 to 0.015, and the relevant-beats-random rate dropped from 0.636 to 0.487 (chance level), confirming the signal tracks proposal relevance rather than masking magnitude alone.
On eight text-only benchmarks (ARC, OpenBookQA, CommonsenseQA, etc.), the Qwen3.5-9B Answer-CED checkpoint showed a mean accuracy change of only -0.27 percentage points, with no single benchmark dropping more than 1.5 points.
Where it can be used
Reward design for RL post-training of VLMs on VQA, counting, or spatial reasoning tasks where hallucination or shortcut reasoning is a concern
Diagnosing or filtering training data to separate 'guessed correctly' answers from 'evidence-grounded' answers when accuracy alone can't distinguish them
Settings needing a low-cost visual grounding audit signal built from weak object-detection boxes (e.g., COCO-style annotations) without question-specific evidence labels
Limits and open work
Experiments are limited to object-centric region interventions using COCO-style boxes; extension to attribute-level, relational, or text-span evidence is proposed conceptually but not separately validated.
Binary yes/no presence-type questions show inherently limited within-group reward variance due to their two-action structure, a reported constraint of the approach.
The counterfactual audit is used only during training; at inference time the trained model runs normally without this check, so whether grounding persists in deployment is a separate question not directly measured here.
Scale and translation perturbations to the proposal region showed no systematic degradation only within the tested range, so robustness beyond that range is unverified.
On blank-placeholder-image ScienceQA items, the CED model correctly refuses to answer ('cannot determine from image') but is scored incorrect against a benchmark that rewards prior-based guessing, showing a scoring mismatch on that subset.
Why it matters
Accuracy-only evaluation can't tell apart a model that guessed correctly from one that actually looked at the image, and this method gives a practical way to reward the difference during training. It matters for anyone building VQA, counting, or spatial-reasoning systems where hallucination or shortcut reasoning from language priors is a known failure mode.
Terms in this paper
VLM (Vision-Language Model) · a model that takes both an image and text as input to produce answers
counterfactual intervention · artificially removing information from a specific region and observing how the model's output changes
GRPO · a reinforcement learning method that samples multiple answers per question and rewards them relative to each other within the group
evidence margin · the difference between how much support drops when the evidence region is masked versus when other regions are masked
Answer-CED vs CoT-CED · two variants of the method — one scores only the final answer span, the other scores the whole reasoning trace
Original abstract (English)
Vision-Language Models (VLMs) should answer from concrete image evidence rather than language priors, dataset shortcuts, or irrelevant visual context. Existing perception-aware post-training methods encourage image use through global perturbations or attention proxies, but they do not test whether a sampled answer causally depends on the local evidence that supports it. We propose Counterfactual Evidence Disentanglement (CED), a training-time evidence audit for VLM grounding. For each response, CED neutralizes an object-centric Evidence Region and compares the resulting support drop against matched non-evidence Regions. We combine this signal with answer correctness inside GRPO, rewarding correct answers that rely on the evidence path rather than shortcut or nuisance paths. CED uses weak object-level proposals, requires no question-specific evidence annotations, and adds no inference-time overhead. Across nine public benchmarks and four backbones, CED outperforms prior RL-based post-training methods, with targeted analyses verifying its object-centric signal.