K-文化的一切——从回归到 K-美妆,发送到您的邮箱订阅邮件

METAL MEDIA

Self- and Other-Labels Induce Bidirectional Bias in LLM Judges

arXiv:2608.180912026-08-20

AI评委打分依据的是标签说的作者是谁,而不是内容本身的好坏

此前人们担心让大模型给自己的输出打分时会偏袒自己,这种自我偏好问题一直难以和文风、真实质量区分开来测量。这项研究改用十个大模型完成一个不带写作风格的叙事元素挑选任务,再用统计方法控制质量和评委严格程度后发现,自我偏好现象基本消失。但当研究者只是给质量相同的两份选择贴上假的自己或他人标签时,无论真实作者是谁,评委都会在自己标签下打高分、在他人标签下打低分。

METAL MEDIA 解读图

AI评委打分依据的是标签说的作者是谁,而不是内容本身的好坏

  1. 01五个商用模型(Claude Opus 4.7、Gemini 3.1 Pro、GPT-5.5、Grok 4.3、Qwen3.6-Plus)和五个开源模型(DeepSeek-V4-Pro、Kimi K2.6、Llama 4 Maverick、Mistral Large 3、Qwen3.6-35B-A3B)共十个大模型,既充当从200个叙事元素池中挑选20个的选择者,也充当给结果打分的评委
  2. 02在不透露作者的盲评实验中,十个评委里有七个给自己的选择打分更高,但用混合效应模型控制了选择质量和评委打分严格度差异后,平均分及四个评分维度中的三个上这种差距消失,只有原创性维度出现反转,评委反而给自己的选择打分更低
  3. 03在标注作者的实验中,研究者事先把质量相当的两份选择配对,再标注为你自己选的或另一个语言模型选的,与真实作者身份无关;结果显示自己标签比他人标签在全部四个评分维度上高出0.29到0.57分,具有统计显著性,而真实作者身份几乎不影响分数
  4. 04十个评委中有九个对标签有反应,其中六个呈对称模式(自己标签打分上升、他人标签打分下降),两个(Claude Opus 4.7、DeepSeek-V4-Pro)只在自己标签下打分上升,一个(GPT-5.5)只在他人标签下打分下降,唯独Llama 4 Maverick在两种标签下都打分下降
这是 METAL MEDIA 制作的解读图,并非论文作者提供的原图。

他们做了什么

  1. 五个商用模型(Claude Opus 4.7、Gemini 3.1 Pro、GPT-5.5、Grok 4.3、Qwen3.6-Plus)和五个开源模型(DeepSeek-V4-Pro、Kimi K2.6、Llama 4 Maverick、Mistral Large 3、Qwen3.6-35B-A3B)共十个大模型,既充当从200个叙事元素池中挑选20个的选择者,也充当给结果打分的评委
  2. 在不透露作者的盲评实验中,十个评委里有七个给自己的选择打分更高,但用混合效应模型控制了选择质量和评委打分严格度差异后,平均分及四个评分维度中的三个上这种差距消失,只有原创性维度出现反转,评委反而给自己的选择打分更低
  3. 在标注作者的实验中,研究者事先把质量相当的两份选择配对,再标注为你自己选的或另一个语言模型选的,与真实作者身份无关;结果显示自己标签比他人标签在全部四个评分维度上高出0.29到0.57分,具有统计显著性,而真实作者身份几乎不影响分数
  4. 十个评委中有九个对标签有反应,其中六个呈对称模式(自己标签打分上升、他人标签打分下降),两个(Claude Opus 4.7、DeepSeek-V4-Pro)只在自己标签下打分上升,一个(GPT-5.5)只在他人标签下打分下降,唯独Llama 4 Maverick在两种标签下都打分下降
Figure 1: Experimental design. Ten LLMs participate in dual roles: as selectors who construct evaluation targets via a constraint selection task, and as judges who evaluate the resulting selections on a 4-axis rubric. Each selection is judged under two conditions: Blind (no source label) and Labeled, where source attribution is experimentally manipulated in a 2×2 design crossing label veracity (TL = true, FL = false) with claimed source (self vs. other).
Figure 1: Experimental design. Ten LLMs participate in dual roles: as selectors who construct evaluation targets via a constraint selection task, and as judges who evaluate the resulting selections on a 4-axis rubric. Each selection is judged under two conditions: Blind (no source label) and Labeled, where source attribution is experimentally manipulated in a 2×2 design crossing label veracity (TL = true, FL = false) with claimed source (self vs. other).
Table 1: 2×2 design for Experiment 2, crossing the actual source of a selection with the label shown to the judge. TL = true label (label matches the actual source); FL = false label (label contradicts the actual source).
Self labelOther label
TL-selfFL-other
Self-producedself selection + self labelself selection + other label
FL-selfTL-other
Other-producedother selection + self labelother selection + other label
Figure 2: Narrative constraint selection task, adapted from Jung et al. (2025). The pool comprises 200 single-sentence constraints from that work, organized into four categories. In each run, an LLM freely selects any 20 constraints from the full pool with no category-level quota. All ten models each perform 30 runs, yielding 300 constraint sets in total, serving as evaluation targets.
Figure 2: Narrative constraint selection task, adapted from Jung et al. (2025). The pool comprises 200 single-sentence constraints from that work, organized into four categories. In each run, an LLM freely selects any 20 constraints from the full pool with no category-level quota. All ten models each perform 30 runs, yielding 300 constraint sets in total, serving as evaluation targets.
Table 2: Raw self-other comparison on the average score, sorted by Δ=Self−Other. Confounds are not controlled; per-dimension results are in Appendix F.
JudgeSelfOtherΔ
Kimi K2.64.883.79+1.09
GPT-5.55.784.92+0.86
DeepSeek-V4-Pro4.864.34+0.52
Claude Opus 4.74.013.50+0.51
Qwen3.6-Plus4.734.49+0.24
Gemini 3.1 Pro4.414.21+0.21
Llama 4 Maverick5.515.47+0.04
Mistral Large 35.946.02−0.08
Grok 4.33.684.35−0.67
Qwen3.6-35B-A3B3.754.65−0.90
Figure 3: Selection patterns by model (300 selections), projected via Multidimensional Scaling (Borg and Groenen, 2005). Each subplot highlights one model’s selections against the full set; models are ordered by within-model Jaccard similarity (descending).
Figure 3: Selection patterns by model (300 selections), projected via Multidimensional Scaling (Borg and Groenen, 2005). Each subplot highlights one model’s selections against the full set; models are ordered by within-model Jaccard similarity (descending).
Table 3: Confound-controlled self-preference coefficient βS from the mixed-effects model, before and after excluding the two low-discrimination judges. Avg. = the average of the four rubric dimensions and the primary outcome; Org. = Originality, Dim. = Dimensionality, Coh. = Coherence, Tel. = Tellability. The positive effects on the average score and three dimensions vanish once low-discrimination judges are removed; only a negative Originality effect persists, which runs counter to self-preference.
Full (10 judges)Excl. low-discrim. (8)
Dim.βS95% CIβS95% CI
Avg.+0.181∗[ 0.13, 0.23 ]+0.003[−0.05, 0.06 ]
Org.−0.144∗[−0.21, −0.08]−0.133∗[−0.21, −0.05]
Dim.+0.316∗[ 0.24, 0.39 ]+0.039[−0.05, 0.12 ]
Coh.+0.276∗[ 0.21, 0.34 ]+0.042[−0.03, 0.12 ]
Tel.+0.278∗[ 0.21, 0.34 ]+0.065[−0.01, 0.14 ]
Figure 4: Score frequency distribution across all Experiment 1 (blind) evaluations, used as a sanity check. Each cell reports the percentage of evaluations in which a judge assigned that score for the given dimension.
Figure 4: Score frequency distribution across all Experiment 1 (blind) evaluations, used as a sanity check. Each cell reports the percentage of evaluations in which a judge assigned that score for the given dimension.
Table 4: Per-condition descriptive statistics for Experiment 2. Mean (SD) across all judges, pairs, and repetitions. TL = label matches actual source; FL = label contradicts actual source.
ActualConditionAverageOrg.Dim.Coh.Tell.
SelfTL-self4.92 (1.04)4.76 (1.24)5.16 (1.55)3.85 (1.74)5.91 (1.09)
-producedFL-other4.50 (1.10)4.28 (1.25)4.60 (1.60)3.56 (1.69)5.55 (1.14)
OtherFL-self4.91 (1.06)4.83 (1.23)5.09 (1.57)3.92 (1.80)5.79 (1.12)
-producedTL-other4.48 (1.16)4.30 (1.19)4.52 (1.68)3.63 (1.82)5.47 (1.22)
Figure 5: Per-judge mean deviation from the Experiment 1 (no-label) baseline under self-labels (red) and other-labels (blue) on the average score, with 95 percent CIs.
Figure 5: Per-judge mean deviation from the Experiment 1 (no-label) baseline under self-labels (red) and other-labels (blue) on the average score, with 95 percent CIs.
Table 5: Fixed-effect estimates from the mixed-effects models for Experiment 2. βL captures the displayed-label effect; βA the actual-source effect (expected ≈0 under successful quality matching); βL​A their interaction. p∗⁣∗∗<.001. Across all dimensions, βL is the only significant fixed effect, and this pattern is robust to excluding two low-discrimination judges (see Appendix H).
Dim.EffectEstimate95% CI𝐩
βL+0.43∗⁣∗∗[+0.39,+0.47]<.001
Avg.βA+0.02[−0.09,+0.14].689
βL​A−0.01[−0.06,+0.05].823
βL+0.53∗⁣∗∗[+0.47,+0.60]<.001
Org.βA−0.03[−0.15,+0.09].663
βL​A−0.05[−0.14,+0.05].346
βL+0.57∗⁣∗∗[+0.50,+0.65]<.001
Dim.βA+0.11[−0.06,+0.27].197
βL​A−0.02[−0.12,+0.08].721
βL+0.29∗⁣∗∗[+0.24,+0.35]<.001
Coh.βA−0.06[−0.26,+0.14].541
βL​A−0.01[−0.09,+0.07].805
βL+0.32∗⁣∗∗[+0.26,+0.38]<.001
Tel.βA+0.07[−0.05,+0.20].257
βL​A+0.05[−0.04,+0.13].261
Figure 9: Average score assigned by each judge (rows) to each selector’s passages (columns). Diagonal cells indicate self-evaluation. Row-level variation reflects judge-level severity differences; column-level variation reflects selection quality. Selectors that receive consistently high scores across judges—notably GPT-5.5 and Kimi K2.6—also score highly on the diagonal, suggesting that the raw self–other gap is driven in part by selection quality rather than genuine self-preference.
Figure 9: Average score assigned by each judge (rows) to each selector’s passages (columns). Diagonal cells indicate self-evaluation. Row-level variation reflects judge-level severity differences; column-level variation reflects selection quality. Selectors that receive consistently high scores across judges—notably GPT-5.5 and Kimi K2.6—also score highly on the diagonal, suggesting that the raw self–other gap is driven in part by selection quality rather than genuine self-preference.
Table 6: Per-judge mean deviation from the Experiment 1 (no-label) baseline under self- and other-label conditions on the average score. Patterns: symmetric (self ↑, other ↓), self-boost (self ↑ only), other-penalty (other ↓ only), general-penalty (both ↓). Rows grouped by pattern; within each group, judges are sorted by the magnitude of the label-induced shift. p∗<.05 (one-sample t-test, H0: μ=0).
JudgeSelf-labelOther-labelPattern
Qwen3.6-35B-A3B+0.131∗−0.523∗symmetric
Gemini 3.1 Pro+0.302∗−0.339∗symmetric
Qwen3.6-Plus+0.175∗−0.444∗symmetric
Kimi K2.6+0.147∗−0.386∗symmetric
Mistral Large 3+0.081∗−0.265∗symmetric
Grok 4.3+0.126∗−0.195∗symmetric
DeepSeek-V4-Pro+0.443∗+0.101self-boost
Claude Opus 4.7+0.275∗−0.023self-boost
GPT-5.5−0.019−0.340∗other-penalty
Llama 4 Maverick−0.192∗−0.384∗general-penalty
Table 7: All models were accessed via OpenRouter222https://openrouter.ai; exact snapshot identifiers are provided. The set spans nine developers and covers frontier general-purpose systems alongside open-weight models of varying scale and training lineage, providing sufficient heterogeneity to probe self-preference across diverse judges and model families. All models are queried with temperature = 1.0 and reasoning_effort = high where each parameter is supported; unsupported parameters were left at their defaults.
ModelOpenRouter ID
Commercial
Claude Opus 4.7anthropic/claude-4.7-opus-20260416
Gemini 3.1 Progoogle/gemini-3.1-pro-preview-20260219
GPT-5.5openai/gpt-5.5-20260423
Grok 4.3x-ai/grok-4.3-20260430
Qwen3.6-Plusqwen/qwen3.6-plus-04-02
Open-weight
DeepSeek-V4-Prodeepseek/deepseek-v4-pro-20260423
Kimi K2.6moonshotai/kimi-k2.6-20260420
Llama 4 Maverickmeta-llama/llama-4-maverick-17b-128e-instruct
Mistral Large 3mistralai/mistral-large-2512
Qwen3.6-35B-A3Bqwen/qwen3.6-35b-a3b-20260415
Table 8: Four-dimension evaluation rubric with anchor descriptions at the endpoints of a 7-point Likert scale. Theoretical grounding for each dimension is provided in the text below.
DimensionAnchor at 1 (Low)Anchor at 7 (High)
OriginalityThe selected constraints lean on familiar story patterns and well-worn tropes, producing a combination that is predictable and conventional.The selected constraints break from familiar story patterns and well-worn tropes, producing a combination that is unpredictable and fresh.
DimensionalityThe selected constraints operate independently; altering one does not affect the meaning or function of the others.The selected constraints are mutually constitutive, each element defining the others’ meaning and weight, producing narrative depth.
CoherenceThe selected constraints contain irreconcilable contradictions that no narrative device can resolve within a single coherent story.The selected constraints integrate seamlessly into a single narrative without logical tensions or unresolved contradictions.
TellabilityThe selected constraints lack a point worth telling—no conflict, question, or stake arises that would make an audience want to hear or tell this story.The selected constraints carry a point worth telling—a clear conflict, urgent question, or meaningful stake that naturally invites narration and sustains attention.
Table 9: Leave-one-out k-NN classification of selection profiles by source model (N=300, chance =0.10, 10 classes). Cohen’s κ ranges from 0.41 to 0.45, indicating moderate chance-corrected agreement. k=5 is used as the primary configuration. p∗⁣∗∗<.001 (binomial test against H0:p=1/10).
kAccuracyCorrect/NMacro-F1κp
10.477143/3000.4270.419***
30.467140/3000.4050.407***
50.507152/3000.4480.452***
70.493148/3000.4240.437***
Table 10: Per-model k-NN classification report (k=5, n=30 per model), sorted by F1 descending. Two asymmetric patterns stand out: Llama 4 Maverick and GPT-5.5 achieve perfect recall but lower precision, whereas Mistral Large 3 achieves perfect precision but is rarely selected as the prediction.
ModelTPFPFNPrec.Rec.F1
Claude Opus 4.7251050.710.830.77
Llama 4 Maverick302200.581.000.73
GPT-5.5303700.451.000.62
Kimi K2.62018100.530.670.59
Qwen3.6-Plus2024100.450.670.54
DeepSeek-V4-Pro76230.540.230.33
Gemini 3.1 Pro812220.400.270.32
Qwen3.6-35B-A3B510250.330.170.22
Grok 4.349260.310.130.19
Mistral Large 330271.000.100.18
Table 11: Within-model selection consistency measured by pairwise Jaccard similarity, computed over all (302)=435 unique pairs from 30 selection runs per model. The random null is the expected mean Jaccard under uniform random selection from 200 constraints (10,000 permutations; mean =0.054, 95% CI: [0.051,0.058]). ∗ denotes p<.0001 relative to the random null (one-tailed permutation test).
ModelMeanSDMinMax
Commercial
Claude Opus 4.70.2180.0860.0260.482
Gemini 3.1 Pro0.1900.0930.0260.600
GPT-5.50.4050.0920.1430.667
Grok 4.30.1200.0600.0000.333
Qwen3.6-Plus0.2350.0790.0260.539
Open-weight
DeepSeek-V4-Pro0.1360.0620.0000.429
Kimi K2.60.2210.0690.0810.429
Llama 4 Maverick0.2470.0910.0530.600
Mistral Large 30.0720.0400.0000.212
Qwen3.6-35B-A3B0.1530.0640.0000.379
Random null0.054
Table 12: Raw self–other comparison per judge across all four rubric dimensions. Δ=Self−Other. Confounds are not controlled.
OriginalityDimensionalityCoherenceTellability
JudgeSelfOtherΔSelfOtherΔSelfOtherΔSelfOtherΔ
Kimi K2.63.434.47−1.045.443.40+2.054.722.44+2.285.934.85+1.08
GPT-5.54.464.74−0.285.995.16+0.835.833.89+1.946.825.87+0.95
DeepSeek-V4-Pro5.685.07+0.604.404.02+0.384.023.48+0.545.334.78+0.56
Claude Opus 4.73.843.87−0.034.133.23+0.912.692.13+0.565.364.77+0.58
Qwen3.6-Plus4.745.00−0.264.964.48+0.473.533.16+0.375.695.30+0.39
Gemini 3.1 Pro3.764.63−0.884.944.34+0.602.862.27+0.596.105.59+0.51
Llama 4 Maverick5.394.89+0.505.895.83+0.064.905.37−0.475.865.78+0.07
Mistral Large 35.985.70+0.286.266.38−0.125.005.41−0.416.526.60−0.07
Grok 4.34.974.42+0.553.024.07−1.042.063.51−1.454.695.41−0.72
Qwen3.6-35B-A3B3.734.62−0.893.964.93−0.972.293.47−1.185.025.59−0.56
Table 13: Robustness check: fixed-effect estimates from the mixed-effects models for Experiment 2 after excluding the two low-discrimination judges (Llama 4 Maverick and Mistral Large 3). Model specification identical to Table 5. βL captures the displayed-label effect; βA the actual-source effect (expected ≈0 under successful quality matching); βL​A their interaction. p∗⁣∗∗<.001. The label effect (βL) remains the only significant fixed effect on every dimension, and its magnitude is comparable to or slightly larger than the full-sample estimate, confirming that the main findings are not driven by the two excluded judges.
Dim.EffectEstimate95% CI𝐩
βL+0.47∗⁣∗∗[+0.43,+0.52]<.001
Avg.βA+0.08[−0.07,+0.22].293
βL​A−0.02[−0.09,+0.05].638
βL+0.57∗⁣∗∗[+0.49,+0.65]<.001
Org.βA−0.12[−0.27,+0.02].089
βL​A−0.03[−0.14,+0.08].595
βL+0.68∗⁣∗∗[+0.59,+0.76]<.001
Dim.βA+0.15[−0.06,+0.35].164
βL​A−0.04[−0.17,+0.08].468
βL+0.27∗⁣∗∗[+0.20,+0.33]<.001
Coh.βA+0.15[−0.09,+0.39].214
βL​A−0.03[−0.12,+0.06].538
βL+0.38∗⁣∗∗[+0.31,+0.45]<.001
Tel.βA+0.13[−0.03,+0.28].101
βL​A+0.04[−0.06,+0.13].462

为什么重要

随着用AI给AI打分的评价方式越来越普及,这项研究说明分数可能仅仅因为标注了谁是作者而改变,与内容真实质量无关,这直接威胁自动化评审结果的可信度。任何使用大模型做评委的场景都应检查评委是否能看到作者或来源信息,因为即使已经排除了文风和质量差异,单凭一个身份标签就足以扭曲评分结果。

本文术语

  • LLM-as-a-judge · 用大语言模型代替人类为输出结果打分或评价,也可用来评价其他AI的输出
  • 自我偏好 · AI评委给自己生成的结果打分更高的倾向
  • 盲评 · 打分时不透露被评内容由谁或哪个模型制作的评估方式
  • 混合效应模型 · 一种能把评委严格程度、内容质量等多种交织影响分离开来分析的统计方法
  • 作者标签 · 打分时标注的声称作者身份,可能与真实制作者不符

论文原文摘要(英文)

As LLM-as-a-judge systems become increasingly widespread, self-preference in LLMs -- the tendency to favor one's own outputs -- raises growing concerns about evaluation reliability. However, it has been studied predominantly on generated text, where stylistic features and response quality are inevitably conflated. As a result, existing measurements cannot separate genuine self-preference from these confounds. We address this by changing the object of evaluation: instead of judging generated text, ten LLMs assess narrative constraint selections, which carry no model-specific stylistic fingerprint yet retain a recoverable model-specific signature. We run two experiments that yield distinct findings. Under blind evaluation, self-preference largely disappears once selection quality and evaluator severity are controlled. It vanishes on three of four rubric dimensions and reverses on the fourth, where judges rate their own selections as less original. Under matched quality, however, self- and other-labels alone -- without naming any model -- shift scores bidirectionally: LLM judges inflate scores for self-labeled selections and deflate those for other-labeled ones regardless of the selection's actual source. We make two contributions: 1) authorship attribution is a distinct driver of evaluation bias, and 2) open-ended, ground-truth-free tasks can serve as controlled instruments for studying LLM judge behavior.

作者 · Songeun Chae, Min Kim, Donghoon Jung, Seojin Choi, Seohyon Jung

在 arXiv 阅读

最新论文

全部论文 →

METAL MEDIA 最新报道

图片来源: Songeun Chae et al., arXiv:2608.18091, CC BY 4.0