Everything K-culture — comebacks to K-beauty, straight to your inboxGet it in your inbox›
Self- and Other-Labels Induce Bidirectional Bias in LLM Judges
arXiv:2608.180912026-08-20
AI judges change their scores based on who they're told the author is, not on actual quality
Prior worry about LLM judges favoring their own outputs, called self-preference, was hard to test cleanly because writing style and actual quality always got mixed together. This study sidesteps that by having ten LLMs pick narrative building blocks rather than write text, then statistically controlling for quality and judge strictness; under this control, self-preference mostly disappeared. But when the researchers simply attached a fake self or other label to identical, quality-matched selections, judges reliably inflated scores under self-labels and deflated them under other-labels regardless of who actually made the selection.
METAL MEDIA explanatory visual
AI judges change their scores based on who they're told the author is, not on actual quality
01Ten LLMs -- five commercial (Claude Opus 4.7, Gemini 3.1 Pro, GPT-5.5, Grok 4.3, Qwen3.6-Plus) and five open-weight (DeepSeek-V4-Pro, Kimi K2.6, Llama 4 Maverick, Mistral Large 3, Qwen3.6-35B-A3B) -- both selected 20 narrative constraints out of a 200-item pool and judged the resulting selections
02In blind evaluation (Experiment 1), 7 of 10 judges rated their own selections higher on average, but once selection quality and each judge's rating severity were controlled with a mixed-effects model, this gap vanished for the average score and three of four rubric dimensions; only Originality showed a reversed effect, with judges scoring their own selections lower
03In labeled evaluation (Experiment 2), quality-matched selection pairs were shown with a label saying 'your own selection' or 'another language model,' independent of true authorship; self-labels raised scores by 0.29-0.57 points over other-labels across all four dimensions, a statistically significant effect, while the actual source made almost no difference
04Nine of ten judges responded to the labels: six showed a symmetric pattern (self-label boosts score, other-label lowers it), two (Claude Opus 4.7, DeepSeek-V4-Pro) only boosted under self-labels, one (GPT-5.5) only penalized under other-labels, and Llama 4 Maverick uniquely lowered scores under both labels
An explanatory diagram made by METAL MEDIA, not a figure supplied by the paper's authors.
What they did
Ten LLMs -- five commercial (Claude Opus 4.7, Gemini 3.1 Pro, GPT-5.5, Grok 4.3, Qwen3.6-Plus) and five open-weight (DeepSeek-V4-Pro, Kimi K2.6, Llama 4 Maverick, Mistral Large 3, Qwen3.6-35B-A3B) -- both selected 20 narrative constraints out of a 200-item pool and judged the resulting selections
In blind evaluation (Experiment 1), 7 of 10 judges rated their own selections higher on average, but once selection quality and each judge's rating severity were controlled with a mixed-effects model, this gap vanished for the average score and three of four rubric dimensions; only Originality showed a reversed effect, with judges scoring their own selections lower
In labeled evaluation (Experiment 2), quality-matched selection pairs were shown with a label saying 'your own selection' or 'another language model,' independent of true authorship; self-labels raised scores by 0.29-0.57 points over other-labels across all four dimensions, a statistically significant effect, while the actual source made almost no difference
Nine of ten judges responded to the labels: six showed a symmetric pattern (self-label boosts score, other-label lowers it), two (Claude Opus 4.7, DeepSeek-V4-Pro) only boosted under self-labels, one (GPT-5.5) only penalized under other-labels, and Llama 4 Maverick uniquely lowered scores under both labels
Figure 1: Experimental design. Ten LLMs participate in dual roles: as selectors who construct evaluation targets via a constraint selection task, and as judges who evaluate the resulting selections on a 4-axis rubric. Each selection is judged under two conditions: Blind (no source label) and Labeled, where source attribution is experimentally manipulated in a 2×2 design crossing label veracity (TL = true, FL = false) with claimed source (self vs. other).
Table 1: 2×2 design for Experiment 2, crossing the actual source of a selection with the label shown to the judge. TL = true label (label matches the actual source); FL = false label (label contradicts the actual source).
Self label
Other label
TL-self
FL-other
Self-produced
self selection + self label
self selection + other label
FL-self
TL-other
Other-produced
other selection + self label
other selection + other label
Figure 2: Narrative constraint selection task, adapted from Jung et al. (2025). The pool comprises 200 single-sentence constraints from that work, organized into four categories. In each run, an LLM freely selects any 20 constraints from the full pool with no category-level quota. All ten models each perform 30 runs, yielding 300 constraint sets in total, serving as evaluation targets.
Table 2: Raw self-other comparison on the average score, sorted by Δ=Self−Other. Confounds are not controlled; per-dimension results are in Appendix F.
Judge
Self
Other
Δ
Kimi K2.6
4.88
3.79
+1.09
GPT-5.5
5.78
4.92
+0.86
DeepSeek-V4-Pro
4.86
4.34
+0.52
Claude Opus 4.7
4.01
3.50
+0.51
Qwen3.6-Plus
4.73
4.49
+0.24
Gemini 3.1 Pro
4.41
4.21
+0.21
Llama 4 Maverick
5.51
5.47
+0.04
Mistral Large 3
5.94
6.02
−0.08
Grok 4.3
3.68
4.35
−0.67
Qwen3.6-35B-A3B
3.75
4.65
−0.90
Figure 3: Selection patterns by model (300 selections), projected via Multidimensional Scaling (Borg and Groenen, 2005). Each subplot highlights one model’s selections against the full set; models are ordered by within-model Jaccard similarity (descending).
Table 3: Confound-controlled self-preference coefficient βS from the mixed-effects model, before and after excluding the two low-discrimination judges. Avg. = the average of the four rubric dimensions and the primary outcome; Org. = Originality, Dim. = Dimensionality, Coh. = Coherence, Tel. = Tellability. The positive effects on the average score and three dimensions vanish once low-discrimination judges are removed; only a negative Originality effect persists, which runs counter to self-preference.
Full (10 judges)
Excl. low-discrim. (8)
Dim.
βS
95% CI
βS
95% CI
Avg.
+0.181∗
[ 0.13, 0.23 ]
+0.003
[−0.05, 0.06 ]
Org.
−0.144∗
[−0.21, −0.08]
−0.133∗
[−0.21, −0.05]
Dim.
+0.316∗
[ 0.24, 0.39 ]
+0.039
[−0.05, 0.12 ]
Coh.
+0.276∗
[ 0.21, 0.34 ]
+0.042
[−0.03, 0.12 ]
Tel.
+0.278∗
[ 0.21, 0.34 ]
+0.065
[−0.01, 0.14 ]
Figure 4: Score frequency distribution across all Experiment 1 (blind) evaluations, used as a sanity check. Each cell reports the percentage of evaluations in which a judge assigned that score for the given dimension.
Table 4: Per-condition descriptive statistics for Experiment 2. Mean (SD) across all judges, pairs, and repetitions. TL = label matches actual source; FL = label contradicts actual source.
Actual
Condition
Average
Org.
Dim.
Coh.
Tell.
Self
TL-self
4.92 (1.04)
4.76 (1.24)
5.16 (1.55)
3.85 (1.74)
5.91 (1.09)
-produced
FL-other
4.50 (1.10)
4.28 (1.25)
4.60 (1.60)
3.56 (1.69)
5.55 (1.14)
Other
FL-self
4.91 (1.06)
4.83 (1.23)
5.09 (1.57)
3.92 (1.80)
5.79 (1.12)
-produced
TL-other
4.48 (1.16)
4.30 (1.19)
4.52 (1.68)
3.63 (1.82)
5.47 (1.22)
Figure 5: Per-judge mean deviation from the Experiment 1 (no-label) baseline under self-labels (red) and other-labels (blue) on the average score, with 95 percent CIs.
Table 5: Fixed-effect estimates from the mixed-effects models for Experiment 2. βL captures the displayed-label effect; βA the actual-source effect (expected ≈0 under successful quality matching); βLA their interaction. p∗∗∗<.001. Across all dimensions, βL is the only significant fixed effect, and this pattern is robust to excluding two low-discrimination judges (see Appendix H).
Dim.
Effect
Estimate
95% CI
𝐩
βL
+0.43∗∗∗
[+0.39,+0.47]
<.001
Avg.
βA
+0.02
[−0.09,+0.14]
.689
βLA
−0.01
[−0.06,+0.05]
.823
βL
+0.53∗∗∗
[+0.47,+0.60]
<.001
Org.
βA
−0.03
[−0.15,+0.09]
.663
βLA
−0.05
[−0.14,+0.05]
.346
βL
+0.57∗∗∗
[+0.50,+0.65]
<.001
Dim.
βA
+0.11
[−0.06,+0.27]
.197
βLA
−0.02
[−0.12,+0.08]
.721
βL
+0.29∗∗∗
[+0.24,+0.35]
<.001
Coh.
βA
−0.06
[−0.26,+0.14]
.541
βLA
−0.01
[−0.09,+0.07]
.805
βL
+0.32∗∗∗
[+0.26,+0.38]
<.001
Tel.
βA
+0.07
[−0.05,+0.20]
.257
βLA
+0.05
[−0.04,+0.13]
.261
Figure 9: Average score assigned by each judge (rows) to each selector’s passages (columns). Diagonal cells indicate self-evaluation. Row-level variation reflects judge-level severity differences; column-level variation reflects selection quality. Selectors that receive consistently high scores across judges—notably GPT-5.5 and Kimi K2.6—also score highly on the diagonal, suggesting that the raw self–other gap is driven in part by selection quality rather than genuine self-preference.
Table 6: Per-judge mean deviation from the Experiment 1 (no-label) baseline under self- and other-label conditions on the average score. Patterns: symmetric (self ↑, other ↓), self-boost (self ↑ only), other-penalty (other ↓ only), general-penalty (both ↓). Rows grouped by pattern; within each group, judges are sorted by the magnitude of the label-induced shift. p∗<.05 (one-sample t-test, H0: μ=0).
Judge
Self-label
Other-label
Pattern
Qwen3.6-35B-A3B
+0.131∗
−0.523∗
symmetric
Gemini 3.1 Pro
+0.302∗
−0.339∗
symmetric
Qwen3.6-Plus
+0.175∗
−0.444∗
symmetric
Kimi K2.6
+0.147∗
−0.386∗
symmetric
Mistral Large 3
+0.081∗
−0.265∗
symmetric
Grok 4.3
+0.126∗
−0.195∗
symmetric
DeepSeek-V4-Pro
+0.443∗
+0.101
self-boost
Claude Opus 4.7
+0.275∗
−0.023
self-boost
GPT-5.5
−0.019
−0.340∗
other-penalty
Llama 4 Maverick
−0.192∗
−0.384∗
general-penalty
Table 7: All models were accessed via OpenRouter222https://openrouter.ai; exact snapshot identifiers are provided. The set spans nine developers and covers frontier general-purpose systems alongside open-weight models of varying scale and training lineage, providing sufficient heterogeneity to probe self-preference across diverse judges and model families. All models are queried with temperature = 1.0 and reasoning_effort = high where each parameter is supported; unsupported parameters were left at their defaults.
Model
OpenRouter ID
Commercial
Claude Opus 4.7
anthropic/claude-4.7-opus-20260416
Gemini 3.1 Pro
google/gemini-3.1-pro-preview-20260219
GPT-5.5
openai/gpt-5.5-20260423
Grok 4.3
x-ai/grok-4.3-20260430
Qwen3.6-Plus
qwen/qwen3.6-plus-04-02
Open-weight
DeepSeek-V4-Pro
deepseek/deepseek-v4-pro-20260423
Kimi K2.6
moonshotai/kimi-k2.6-20260420
Llama 4 Maverick
meta-llama/llama-4-maverick-17b-128e-instruct
Mistral Large 3
mistralai/mistral-large-2512
Qwen3.6-35B-A3B
qwen/qwen3.6-35b-a3b-20260415
Table 8: Four-dimension evaluation rubric with anchor descriptions at the endpoints of a 7-point Likert scale. Theoretical grounding for each dimension is provided in the text below.
Dimension
Anchor at 1 (Low)
Anchor at 7 (High)
Originality
The selected constraints lean on familiar story patterns and well-worn tropes, producing a combination that is predictable and conventional.
The selected constraints break from familiar story patterns and well-worn tropes, producing a combination that is unpredictable and fresh.
Dimensionality
The selected constraints operate independently; altering one does not affect the meaning or function of the others.
The selected constraints are mutually constitutive, each element defining the others’ meaning and weight, producing narrative depth.
Coherence
The selected constraints contain irreconcilable contradictions that no narrative device can resolve within a single coherent story.
The selected constraints integrate seamlessly into a single narrative without logical tensions or unresolved contradictions.
Tellability
The selected constraints lack a point worth telling—no conflict, question, or stake arises that would make an audience want to hear or tell this story.
The selected constraints carry a point worth telling—a clear conflict, urgent question, or meaningful stake that naturally invites narration and sustains attention.
Table 9: Leave-one-out k-NN classification of selection profiles by source model (N=300, chance =0.10, 10 classes). Cohen’s κ ranges from 0.41 to 0.45, indicating moderate chance-corrected agreement. k=5 is used as the primary configuration. p∗∗∗<.001 (binomial test against H0:p=1/10).
k
Accuracy
Correct/N
Macro-F1
κ
p
1
0.477
143/300
0.427
0.419
***
3
0.467
140/300
0.405
0.407
***
5
0.507
152/300
0.448
0.452
***
7
0.493
148/300
0.424
0.437
***
Table 10: Per-model k-NN classification report (k=5, n=30 per model), sorted by F1 descending. Two asymmetric patterns stand out: Llama 4 Maverick and GPT-5.5 achieve perfect recall but lower precision, whereas Mistral Large 3 achieves perfect precision but is rarely selected as the prediction.
Model
TP
FP
FN
Prec.
Rec.
F1
Claude Opus 4.7
25
10
5
0.71
0.83
0.77
Llama 4 Maverick
30
22
0
0.58
1.00
0.73
GPT-5.5
30
37
0
0.45
1.00
0.62
Kimi K2.6
20
18
10
0.53
0.67
0.59
Qwen3.6-Plus
20
24
10
0.45
0.67
0.54
DeepSeek-V4-Pro
7
6
23
0.54
0.23
0.33
Gemini 3.1 Pro
8
12
22
0.40
0.27
0.32
Qwen3.6-35B-A3B
5
10
25
0.33
0.17
0.22
Grok 4.3
4
9
26
0.31
0.13
0.19
Mistral Large 3
3
0
27
1.00
0.10
0.18
Table 11: Within-model selection consistency measured by pairwise Jaccard similarity, computed over all (302)=435 unique pairs from 30 selection runs per model. The random null is the expected mean Jaccard under uniform random selection from 200 constraints (10,000 permutations; mean =0.054, 95% CI: [0.051,0.058]). ∗ denotes p<.0001 relative to the random null (one-tailed permutation test).
Model
Mean
SD
Min
Max
Commercial
Claude Opus 4.7
0.218
0.086
0.026
0.482
∗
Gemini 3.1 Pro
0.190
0.093
0.026
0.600
∗
GPT-5.5
0.405
0.092
0.143
0.667
∗
Grok 4.3
0.120
0.060
0.000
0.333
∗
Qwen3.6-Plus
0.235
0.079
0.026
0.539
∗
Open-weight
DeepSeek-V4-Pro
0.136
0.062
0.000
0.429
∗
Kimi K2.6
0.221
0.069
0.081
0.429
∗
Llama 4 Maverick
0.247
0.091
0.053
0.600
∗
Mistral Large 3
0.072
0.040
0.000
0.212
∗
Qwen3.6-35B-A3B
0.153
0.064
0.000
0.379
∗
Random null
0.054
—
Table 12: Raw self–other comparison per judge across all four rubric dimensions. Δ=Self−Other. Confounds are not controlled.
Originality
Dimensionality
Coherence
Tellability
Judge
Self
Other
Δ
Self
Other
Δ
Self
Other
Δ
Self
Other
Δ
Kimi K2.6
3.43
4.47
−1.04
5.44
3.40
+2.05
4.72
2.44
+2.28
5.93
4.85
+1.08
GPT-5.5
4.46
4.74
−0.28
5.99
5.16
+0.83
5.83
3.89
+1.94
6.82
5.87
+0.95
DeepSeek-V4-Pro
5.68
5.07
+0.60
4.40
4.02
+0.38
4.02
3.48
+0.54
5.33
4.78
+0.56
Claude Opus 4.7
3.84
3.87
−0.03
4.13
3.23
+0.91
2.69
2.13
+0.56
5.36
4.77
+0.58
Qwen3.6-Plus
4.74
5.00
−0.26
4.96
4.48
+0.47
3.53
3.16
+0.37
5.69
5.30
+0.39
Gemini 3.1 Pro
3.76
4.63
−0.88
4.94
4.34
+0.60
2.86
2.27
+0.59
6.10
5.59
+0.51
Llama 4 Maverick
5.39
4.89
+0.50
5.89
5.83
+0.06
4.90
5.37
−0.47
5.86
5.78
+0.07
Mistral Large 3
5.98
5.70
+0.28
6.26
6.38
−0.12
5.00
5.41
−0.41
6.52
6.60
−0.07
Grok 4.3
4.97
4.42
+0.55
3.02
4.07
−1.04
2.06
3.51
−1.45
4.69
5.41
−0.72
Qwen3.6-35B-A3B
3.73
4.62
−0.89
3.96
4.93
−0.97
2.29
3.47
−1.18
5.02
5.59
−0.56
Table 13: Robustness check: fixed-effect estimates from the mixed-effects models for Experiment 2 after excluding the two low-discrimination judges (Llama 4 Maverick and Mistral Large 3). Model specification identical to Table 5. βL captures the displayed-label effect; βA the actual-source effect (expected ≈0 under successful quality matching); βLA their interaction. p∗∗∗<.001. The label effect (βL) remains the only significant fixed effect on every dimension, and its magnitude is comparable to or slightly larger than the full-sample estimate, confirming that the main findings are not driven by the two excluded judges.
Dim.
Effect
Estimate
95% CI
𝐩
βL
+0.47∗∗∗
[+0.43,+0.52]
<.001
Avg.
βA
+0.08
[−0.07,+0.22]
.293
βLA
−0.02
[−0.09,+0.05]
.638
βL
+0.57∗∗∗
[+0.49,+0.65]
<.001
Org.
βA
−0.12
[−0.27,+0.02]
.089
βLA
−0.03
[−0.14,+0.08]
.595
βL
+0.68∗∗∗
[+0.59,+0.76]
<.001
Dim.
βA
+0.15
[−0.06,+0.35]
.164
βLA
−0.04
[−0.17,+0.08]
.468
βL
+0.27∗∗∗
[+0.20,+0.33]
<.001
Coh.
βA
+0.15
[−0.09,+0.39]
.214
βLA
−0.03
[−0.12,+0.06]
.538
βL
+0.38∗∗∗
[+0.31,+0.45]
<.001
Tel.
βA
+0.13
[−0.03,+0.28]
.101
βLA
+0.04
[−0.06,+0.13]
.462
Why it matters
As AI-graded evaluation becomes standard practice, this shows that scores can shift purely from an authorship tag rather than actual content quality, which is a direct threat to trusting automated judging pipelines. It suggests that anyone deploying LLM judges should check whether authorship or source information is exposed to the judge, since that alone can skew results even after style and quality confounds are removed.
Terms in this paper
LLM-as-a-judge · Using a large language model instead of a human to score or evaluate outputs, including other AI outputs
self-preference · The tendency of an AI judge to rate its own outputs more favorably than others'
blind evaluation · Scoring done without revealing who or what produced the item being judged
mixed-effects model · A statistical method that separates out overlapping influences like judge strictness and content quality
authorship label · A tag shown at judging time claiming who made the item, which may or may not match the true creator
Original abstract (English)
As LLM-as-a-judge systems become increasingly widespread, self-preference in LLMs -- the tendency to favor one's own outputs -- raises growing concerns about evaluation reliability. However, it has been studied predominantly on generated text, where stylistic features and response quality are inevitably conflated. As a result, existing measurements cannot separate genuine self-preference from these confounds. We address this by changing the object of evaluation: instead of judging generated text, ten LLMs assess narrative constraint selections, which carry no model-specific stylistic fingerprint yet retain a recoverable model-specific signature. We run two experiments that yield distinct findings. Under blind evaluation, self-preference largely disappears once selection quality and evaluator severity are controlled. It vanishes on three of four rubric dimensions and reverses on the fourth, where judges rate their own selections as less original. Under matched quality, however, self- and other-labels alone -- without naming any model -- shift scores bidirectionally: LLM judges inflate scores for self-labeled selections and deflate those for other-labeled ones regardless of the selection's actual source. We make two contributions: 1) authorship attribution is a distinct driver of evaluation bias, and 2) open-ended, ground-truth-free tasks can serve as controlled instruments for studying LLM judge behavior.
Authors · Songeun Chae, Min Kim, Donghoon Jung, Seojin Choi, Seohyon Jung