Figure 1: Left: Task setup with a counting query over two views of the same scene. Upper right: A pretrained VLM gives the same prior-based answer for both views; Evidence-RL inspects the visual evidence and adjusts its answer when the scene changes. Lower right: Causal path decomposition: a candidate answer can be supported by the target visual evidence, by irrelevant context, or by language priors, which motivates CED to test whether the response depends specifically on the evidence path.
Table 1: Signal-validation statistics on counting and presence tasks.
Task
Non-constant raw reward
Zero variance
Mean reward std
Same answer, different reward
Mean negation
Counting
99.5%
24.0%
0.1299
90.0%
0.003
Presence
21.1%
79.0%
0.0497
5.0%
0.777
Figure 2: Overview of Evidence-RL during post-training. A policy samples candidate answers, and CED scores each answer by applying the same feature-space intervention to the proposed Evidence Region and to matched non-evidence Regions. The contrastive support drop yields an evidence margin, which is combined with correctness to form the GRPO reward. The counterfactual audit is used only during training; inference uses the trained VLM normally.
Table 2: Nine-benchmark evaluation using Qwen2.5-VL-7B as the base model for Ours. Backbone, baseline, and benchmark sources are cited in the setup paragraph. Best bold; second-best underlined. Each Δ row reports absolute gain over the corresponding base.
Grounding
General Reasoning
Model
CountBench
SpatialEval
Hallusion
VLMs AreBlind
FREAK
MathVista
MMBench
MMMU
ScienceQA
Avg
VLM backbones
LLaVA-v1.6-7B
55.60
23.60
51.00
28.23
9.50
22.80
71.30
29.00
57.20
38.69
InternVL3.5-8B
86.90
24.00
69.60
54.92
12.80
51.90
83.40
50.80
89.90
58.25
Qwen2.5-VL-3B
70.71
53.94
65.37
42.46
11.14
52.10
80.40
50.29
74.93
55.70
Qwen2.5-VL-7B
81.82
59.21
68.64
46.44
12.18
63.10
84.73
50.63
83.98
61.19
RL-based methods
VAPO-Thinker-7B
86.90
60.80
71.20
48.62
13.10
51.10
81.80
44.60
82.30
60.05
Δ
+5.10
+1.59
+2.60
+2.18
+0.92
−12.00
−2.90
−6.00
−1.70
−1.14
VPPO-7B
85.90
61.80
68.20
49.35
12.70
67.90
84.90
52.10
88.40
63.47
Δ
+4.10
+2.59
−0.40
+2.91
+0.52
+4.80
+0.20
+1.50
+4.40
+2.28
Perception-R1-7B
84.90
59.70
66.70
45.22
12.20
67.10
81.70
48.10
82.30
60.88
Δ
+3.10
+0.49
−1.90
−1.22
+0.02
+4.00
−3.00
−2.50
−1.70
−0.31
PAPO-G-H
90.90
64.80
69.50
47.92
12.80
69.40
82.80
50.50
85.60
63.80
Δ
+9.10
+5.59
+0.90
+1.48
+0.62
+6.30
−1.90
−0.10
+1.60
+2.61
VLM-R1
74.80
56.40
66.60
44.15
11.00
60.30
79.90
48.60
73.60
57.26
Δ vs. Qwen2.5-VL-3B
+4.10
+2.50
+1.20
+1.69
−0.14
+8.20
−0.50
−1.70
−1.30
+1.56
SophiaVL-R1
82.83
61.38
66.87
47.09
24.18
66.30
87.41
52.18
87.64
63.99
Δ
+1.01
+2.17
−1.77
+0.65
+12.00
+3.20
+2.68
+1.55
+3.66
+2.80
Ours
88.89
63.34
70.08
54.02
26.40
69.40
87.53
57.47
87.08
67.13
Δ
+7.07
+4.13
+1.44
+7.58
+14.22
+6.30
+2.80
+6.84
+3.10
+5.94
Figure 3: A case showing why correctness-only reward is insufficient. Two rollouts from the same GRPO group both answer “4” correctly, but the CED gate g(m) creates a 7× reward gap between a prior-based rollout (g(m)=0.18, R=0.11) and an evidence-grounded rollout (g(m)=1.00, R=0.78), enabling GRPO to select the grounded trajectory.
Table 3: Controlled comparison on Qwen3.5-9B with matched data, router, and compute. Δ is the mean gain over the frozen base; ΔCED is Answer-CED minus correctness-only.
Grounding
General Reasoning
Method
CountBench
SpatialEval
Hallusion
VLMs AreBlind
FREAK
MathVista
MMBench
MMMU
ScienceQA
Avg
Δ
λ=0 (corr.-only)
89.90
40.45
57.93
72.90
18.73
73.40
79.44
36.00
82.24
61.22
−1.24
VPPO (retrained)
86.87
38.32
58.37
72.97
19.07
72.50
79.62
37.33
81.09
60.68
−1.78
PAPO (retrained)†
86.87
8.59
58.46
69.96
19.40
14.00
44.87
33.22
34.24
41.07
−21.39
Answer-CED
93.90
54.35
78.10
73.10
23.12
83.75
88.58
75.67
93.61
73.80
+11.34
ΔCED
+4.00
+13.90
+20.17
+0.20
+4.39
+10.35
+9.14
+39.67
+11.37
+12.58
–
Figure 4: Training dynamics for Answer-CED and CoT-CED on Qwen3.5-9B. CoT-CED produces larger training-side rewards, but the downstream comparison in Table 5 favors Answer-CED on average; we therefore use Answer-CED as the default variant.
Table 4: Cross-backbone validation of Ours. Each block reports the frozen base, Ours, and absolute Δ against that block’s base. The Qwen3.5-9B block additionally reports 3-seed mean±std, with its mean Δ in the last column.
Grounding
General Reasoning
Model
CountBench
SpatialEval
Hallusion
VLMs AreBlind
FREAK
MathVista
MMBench
MMMU
ScienceQA
Avg. Δ
Backbone: Qwen2.5-VL-3B
Qwen2.5-VL-3B
70.71
53.94
65.37
42.46
11.14
52.10
80.40
50.29
74.93
–
Ours
70.71
55.56
67.40
49.51
25.51
60.40
82.55
50.57
80.43
–
Δ
+0.00
+1.62
+2.03
+7.05
+14.37
+8.30
+2.15
+0.28
+5.50
+4.59
Backbone: Qwen2.5-VL-7B
Qwen2.5-VL-7B
81.82
59.21
68.64
46.44
12.18
63.10
84.73
50.63
83.98
–
Ours
88.89
63.34
70.08
54.02
26.40
69.40
87.53
57.47
87.08
–
Δ
+7.07
+4.13
+1.44
+7.58
+14.22
+6.30
+2.80
+6.84
+3.10
+5.94
Backbone: Qwen3-VL-8B-Instruct
Qwen3-VL-8B
94.90
65.26
72.86
67.01
28.85
67.54
89.15
57.82
91.99
–
Ours
95.92
67.75
74.30
68.88
31.18
68.04
89.58
58.85
92.82
–
Δ
+1.02
+2.49
+1.44
+1.87
+2.33
+0.50
+0.43
+1.03
+0.83
+1.33
Backbone: Qwen3.5-9B
Qwen3.5-9B
80.80
52.73
47.20
46.44
9.95
80.00
88.03
66.21
90.78
–
Ours
93.90
54.35
78.10
73.10
23.12
83.75
88.58
75.67
93.61
–
Δ
+13.10
+1.62
+30.90
+26.66
+13.17
+3.75
+0.55
+9.46
+2.83
+11.34
Ours (3 seeds)
93.60 ±0.48
56.04 ±0.27
78.35 ±0.21
73.37 ±0.17
23.11 ±0.40
84.40 ±0.50
89.87 ±0.04
75.39 ±0.22
94.46 ±0.13
+11.83
Figure 5: CED probe robustness under proposal perturbation. (a) Replacing the COCO proposal with a random object box as Ωev collapses the evidence margin to near zero across all seven task types, confirming that the signal tracks proposal relevance rather than masking magnitude. (b) IoU-graded degradation on the count-exclusion diagnostic: m¯ decreases monotonically as spatial overlap with the original proposal drops, approaching the random-region floor at lowest IoU.
Table 5: Answer vs. CoT variants on Qwen3.5-9B; Avg. Δ is the mean gain over the base.
Grounding
General Reasoning
Variant
CountBench
SpatialEval
Hallusion
VLMs AreBlind
FREAK
MathVista
MMBench
MMMU
ScienceQA
Avg. Δ
Answer
93.90
54.35
78.10
73.10
23.12
83.75
88.58
75.67
93.61
+11.34
CoT
92.90
53.80
78.10
69.09
30.35
81.88
88.53
68.23
94.66
+10.60
Figure A.1: CED probe robustness under proposal perturbation on the count-exclusion diagnostic. (a) Mean evidence margin m¯ across IoU bins, with the random-region baseline as a grey dashed reference. (b,c) m¯ under scale and translation perturbations, with the original ±1 SE band in red. Error bars are ±1 SE.
Table A.1: Complete hyperparameter listing. “Paper notation” gives the corresponding symbol in the main text when applicable.
Category
Parameter
Value
Paper notation
RL & optimization
Learning rate
lr
1×10−5
—
KL penalty coefficient
kl_coeff
0.01
—
GRPO group size
group_size
32
—
Training steps
n_steps
2,000
—
Trainable layers
n_trainable_layers
4 (last)
—
Precision
dtype
bfloat16
—
Generation & sampling
Temperature
temperature
1.0
—
Top-p
top_p
0.95
—
Max new tokens (evidence)
max_new_tokens
32
—
Max response tokens (full)
max_response_tokens
64
—
Prompt format
short_evidence_v1
—
—
Reward function (Equation 4)
Response reward weight
alpha_resp
0.70
see note†
Answer reward weight
alpha_ans
0.30
see note†
Response logprob temp.
tau_resp
0.20
—
Answer logprob temp.
tau_ans
1.00
—
Gate temperature
tau_resp (shared)
0.20
τg
Gate baseline at m=0
(derived)
0.50
g(0)
Evidence tie-breaker
evidence_eps
0.10
εtie
Additive evidence weight
evidence_eps (shared)
0.10
λ
Reward clipping
[min, max]
[−1.25, 1.00]
—
Format penalty
—
−0.75
—
Data & evaluation
Rerank candidates
rerank_num_candidates
6
—
non-evidence Regions per sample
K
3
K
Figure A.2: Qualitative examples of same-answer-different-reward under the CED method. All rollouts share the same answer, while CED produces reward spans of 0.50–0.67 according to visual evidence quality.
Table A.2: Methodological comparison with perception-aware VLM RL methods. The Causal audit column indicates whether the method tests causal evidence dependence for a candidate answer.
Method
Signal source
Granularity
Causal audit
Stage
PAPO [36]
Global mask KL
Global
No
Training
VPPO [10]
Attention weights
Token-level
No
Training
Vision-SR1†
Text self-verification
Description-level
No
Training
PaLMR‡
LLM-as-Judge
Trajectory-level
No
Training
SophiaVL-R1 [6]
Thinking-reward RM
Trajectory-level
No
Training
PEARL§
Perception checklist
Sample-level
No
Training
VAPO¶
Trajectory anchoring
Token-level
No
Training
VLM-R1*
Self-verification
Response-level
No
Training
Perception-R1††
Proxy localization
Region-levela
No
Training
CED (Ours)
Counterfactual intervention
Region-level
Yes
Training
Figure A.3: Within-group zero-variance rate across task families. Evidence-intensive tasks show a large reduction from correctness-only reward to our reward; binary-dominated groups remain near high zero-variance because of their two-action structure.
Table A.3: Cross-model evidence sensitivity validation (N=50 per task per model).
Model
Task
Relevant > Random rate
Mean margin
InternVL3.5-8B
Counting
0.68
0.0070
Presence
0.62
0.0737
LLaVA-v1.6-7B
Counting
0.78
0.0074
Presence
0.60
0.1625
Figure A.4: Failure cases on presence tasks: identical negation answers with identical rewards yield zero within-group variance.
Table A.4: Blindfold test results (N=500 per task family).
Condition
Task
Rel. > Rand. rate
Mean Delta
Normal
Counting
0.634
+1.928
Normal
Presence
0.550
+0.780
Blindfold
Counting
0.404
−0.114
Blindfold
Presence
0.448
+0.070
Table A.5: Perturbation-family ablation within the local intervention family, evaluated by logits-JS AUC.
Key mode
Mean repl.
Zero repl.
Gaussian-noise repl.
prompt_last
0.6780
0.6602
0.6299
answer_first
0.6592
0.6215
0.6273
Average
0.6686
0.6409
0.6286
Table A.6: Gaussian-noise profile under prompt_last. Stability drops under logits_only.
Config
Fixed metric
Best metric
Best paired AUC
logits24
0.6299
0.6299
0.7108
logits_only
0.5753
0.5920
0.6407
Table A.7: Evidence weight (λ) sensitivity. λ=0.00 is correctness-only.
λ
Non-const. rate
Reward std
Zero-var. rate
0.00
100%
0.0528
0%
0.05
100%
0.0508
0%
0.10
96%
0.0597
4%
0.15
100%
0.0415
4%
0.20
100%
0.0491
0%
Table A.8: Gate temperature (τg) sensitivity.
τg
Non-const. rate
Reward std
Zero-var. rate
0.10
100%
0.0568
4%
0.20
96%
0.0597
4%
0.30
96%
0.0469
4%
0.50
100%
0.0378
0%
Table A.9: Proposal-perturbation results on count-exclusion (n=41). Higher values indicate stronger evidence dependence; the random row gives the floor without spatial targeting.
Axis
Condition
m¯
Δ
Rel. rate
IoU stratification
[0.7, 1.0]
0.501
2.032
0.732
[0.5, 0.7)
0.379
1.282
0.683
[0.3, 0.5)
0.341
1.386
0.683
[0.0, 0.3)
0.219
0.555
0.561
Scale
0.50×
0.377
1.113
0.732
0.75×
0.481
1.750
0.732
1.25×
0.433
2.230
0.683
1.50×
0.295
2.003
0.634
2.00×
0.382
1.880
0.707
Shift
5%
0.304
1.602
0.634
10%
0.428
1.457
0.707
20%
0.532
1.899
0.732
original proposal (reference)
0.386
1.852
0.683
random region (floor)
0.147
0.277
0.537
Table A.10: Unified-condition baseline comparison on the offline rerank dataset (N=1,353).
Method
Reward mean
Signal strength
Signal-corr
Vanilla GRPO
−0.149
N/A
N/A
PAPO-Lite
−0.156
KL =0.024
0.034
CED (Ours)
+0.448
Margin =−0.007
Selective
Table A.11: Unified-condition method comparison on the counting shared-candidate slice: 94 samples with at least two distinct candidates.
Method
Acc.
Strong inf.
rnc↑
flip ↑
LogProb
94.7
75.0
–
–
Correctness-only
97.9
100.0
12.8
4.3
CED (Ours)
97.9
100.0
100.0
76.6
Perception-R1-Lite
97.9
100.0
13.8
5.3
VLM-R1-Lite∗
94.7
75.0
0.0
0.0
VPPO-Lite∗
94.7
75.0
0.0
0.0
PAPO-Lite∗
94.7
75.0
0.0
0.0
∗ No reward-based reranking recovered; falls back to log probability.
Table A.12: CoT-CED vs. Answer-CED: training-end summary over 2,000 steps.
Metric
CoT-CED
Answer-CED
Mean training reward
+0.259
−0.009
Final training correctness
43.4%
34.2%
Table A.13: End-of-training CoT behavior on held-out prompts. CoT-CED collapses toward an object-cue answer format, while Answer-CED preserves natural multi-sentence reasoning.
Metric
Benchmark
CoT-CED
Answer-CED
Mean CoT length (tokens)
Counting
3.6
49.3
Median CoT length (tokens)
Counting
3
47
p90 CoT length (tokens)
Counting
6
78
Mean CoT length (tokens)
HallusionBench
3.8
82.1
Median CoT length (tokens)
HallusionBench
3
61
p90 CoT length (tokens)
HallusionBench
6
173
Image-referential opener rate
Counting
1.0%
79.8%
Image-referential opener rate
HallusionBench
5.2%
64.2%
Unique-token ratio
Counting
99.1%
75.8%
Unique-token ratio
HallusionBench
96.8%
68.1%
Table A.14: Representative side-by-side generations. Both variants produce the correct final answer, but CoT-CED reduces the chain to an object cue, while Answer-CED keeps an explicit visual check.
Prompt
CoT-CED
Answer-CED
How many beds are there in the image?
The beds / 2
Got it, let’s look at the image. There’s a loft bed with a lower bed. The upper bunk is one, the lower bed is another. So that’s two beds.
How many weights are there in the image?
The weights / Two
Got it, let’s look at the image. The man is holding two dumbbells, one in each hand. There’s one in the left hand and one in the right hand. So that’s two weights.
Vision-Language Models (VLMs) should answer from concrete image evidence rather than language priors, dataset shortcuts, or irrelevant visual context. Existing perception-aware post-training methods encourage image use through global perturbations or attention proxies, but they do not test whether a sampled answer causally depends on the local evidence that supports it. We propose Counterfactual Evidence Disentanglement (CED), a training-time evidence audit for VLM grounding. For each response, CED neutralizes an object-centric Evidence Region and compares the resulting support drop against matched non-evidence Regions. We combine this signal with answer correctness inside GRPO, rewarding correct answers that rely on the evidence path rather than shortcut or nuisance paths. CED uses weak object-level proposals, requires no question-specific evidence annotations, and adds no inference-time overhead. Across nine public benchmarks and four backbones, CED outperforms prior RL-based post-training methods, with targeted analyses verifying its object-centric signal.