컴백부터 K-뷰티까지 — K-컬쳐의 모든 것을 메일로 받아보세요메일로 받아보기

METAL MEDIA

AI 심사위원은 글의 질이 아니라 '누가 썼다는 말표'만으로 점수를 바꾼다

arXiv:2608.180912026-08-20

Self- and Other-Labels Induce Bidirectional Bias in LLM Judges

AI 심사위원은 글의 질이 아니라 '누가 썼다는 말표'만으로 점수를 바꾼다

LLM에게 자기 답을 채점하게 하면 자기 것을 더 후하게 평가한다는 '자기선호' 문제가 있었는데, 이 연구는 글쓰기 스타일과 실제 품질을 완전히 분리한 실험으로 이를 다시 검증했다. 문체가 없는 '이야기 구성요소 선택' 과제를 쓰고 품질과 채점자 성향을 통계로 통제하자 자기선호는 대부분 사라졌다. 하지만 실제 저자를 숨기고 '이건 당신이 고른 것' 또는 '다른 AI가 고른 것'이라는 라벨만 붙였을 때는, 진짜 저자와 무관하게 자기 라벨엔 점수를 올리고 타인 라벨엔 점수를 내리는 현상이 뚜렷하게 나타났다.

METAL MEDIA 해설 도표

AI 심사위원은 글의 질이 아니라 '누가 썼다는 말표'만으로 점수를 바꾼다

  1. 01Claude Opus 4.7, Gemini 3.1 Pro, GPT-5.5, Grok 4.3, Qwen3.6-Plus 등 상업 모델 5개와 DeepSeek-V4-Pro, Kimi K2.6, Llama 4 Maverick, Mistral Large 3, Qwen3.6-35B-A3B 등 오픈소스 모델 5개, 총 10개 LLM이 '이야기 구성요소 200개 중 20개 고르기' 과제의 선택자이자 채점자 역할을 모두 맡았다
  2. 02실험1(블라인드 채점)에서는 원작자를 감춘 채 점수를 매겼는데, 10개 중 7개 모델이 자기 선택에 더 높은 점수를 줬지만, 선택자의 실제 품질과 채점자별 엄격함 차이를 통계 모델로 통제하자 독창성 항목만 자기 것에 오히려 낮은 점수를 주는 역전 현상이 남고 나머지는 유의미한 차이가 사라졌다
  3. 03실험2(라벨 채점)에서는 품질을 사전에 맞춘 두 선택지에 '당신이 고른 것' 또는 '다른 AI가 고른 것'이라는 라벨만 붙여 실제 출처와 무관하게 조작했는데, 라벨이 자기(self)면 점수가 오르고 타인(other)이면 내려가는 현상이 4개 평가 항목 전부에서 0.29~0.57점 차이로 통계적으로 유의하게 나타났으며 실제 출처 효과는 거의 없었다
  4. 0410개 채점자 중 9개가 라벨에 반응했고 그중 6개는 자기 라벨엔 올리고 타인 라벨엔 내리는 양방향 패턴, 2개(Claude Opus 4.7, DeepSeek-V4-Pro)는 자기 라벨에만 반응, 1개(GPT-5.5)는 타인 라벨에만 반응, Llama 4 Maverick만 양쪽 다 낮추는 예외 패턴을 보였다
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.

무엇을 했나

  1. Claude Opus 4.7, Gemini 3.1 Pro, GPT-5.5, Grok 4.3, Qwen3.6-Plus 등 상업 모델 5개와 DeepSeek-V4-Pro, Kimi K2.6, Llama 4 Maverick, Mistral Large 3, Qwen3.6-35B-A3B 등 오픈소스 모델 5개, 총 10개 LLM이 '이야기 구성요소 200개 중 20개 고르기' 과제의 선택자이자 채점자 역할을 모두 맡았다
  2. 실험1(블라인드 채점)에서는 원작자를 감춘 채 점수를 매겼는데, 10개 중 7개 모델이 자기 선택에 더 높은 점수를 줬지만, 선택자의 실제 품질과 채점자별 엄격함 차이를 통계 모델로 통제하자 독창성 항목만 자기 것에 오히려 낮은 점수를 주는 역전 현상이 남고 나머지는 유의미한 차이가 사라졌다
  3. 실험2(라벨 채점)에서는 품질을 사전에 맞춘 두 선택지에 '당신이 고른 것' 또는 '다른 AI가 고른 것'이라는 라벨만 붙여 실제 출처와 무관하게 조작했는데, 라벨이 자기(self)면 점수가 오르고 타인(other)이면 내려가는 현상이 4개 평가 항목 전부에서 0.29~0.57점 차이로 통계적으로 유의하게 나타났으며 실제 출처 효과는 거의 없었다
  4. 10개 채점자 중 9개가 라벨에 반응했고 그중 6개는 자기 라벨엔 올리고 타인 라벨엔 내리는 양방향 패턴, 2개(Claude Opus 4.7, DeepSeek-V4-Pro)는 자기 라벨에만 반응, 1개(GPT-5.5)는 타인 라벨에만 반응, Llama 4 Maverick만 양쪽 다 낮추는 예외 패턴을 보였다
Figure 1: Experimental design. Ten LLMs participate in dual roles: as selectors who construct evaluation targets via a constraint selection task, and as judges who evaluate the resulting selections on a 4-axis rubric. Each selection is judged under two conditions: Blind (no source label) and Labeled, where source attribution is experimentally manipulated in a 2×2 design crossing label veracity (TL = true, FL = false) with claimed source (self vs. other).
Figure 1: Experimental design. Ten LLMs participate in dual roles: as selectors who construct evaluation targets via a constraint selection task, and as judges who evaluate the resulting selections on a 4-axis rubric. Each selection is judged under two conditions: Blind (no source label) and Labeled, where source attribution is experimentally manipulated in a 2×2 design crossing label veracity (TL = true, FL = false) with claimed source (self vs. other).
Table 1: 2×2 design for Experiment 2, crossing the actual source of a selection with the label shown to the judge. TL = true label (label matches the actual source); FL = false label (label contradicts the actual source).
Self labelOther label
TL-selfFL-other
Self-producedself selection + self labelself selection + other label
FL-selfTL-other
Other-producedother selection + self labelother selection + other label
Figure 2: Narrative constraint selection task, adapted from Jung et al. (2025). The pool comprises 200 single-sentence constraints from that work, organized into four categories. In each run, an LLM freely selects any 20 constraints from the full pool with no category-level quota. All ten models each perform 30 runs, yielding 300 constraint sets in total, serving as evaluation targets.
Figure 2: Narrative constraint selection task, adapted from Jung et al. (2025). The pool comprises 200 single-sentence constraints from that work, organized into four categories. In each run, an LLM freely selects any 20 constraints from the full pool with no category-level quota. All ten models each perform 30 runs, yielding 300 constraint sets in total, serving as evaluation targets.
Table 2: Raw self-other comparison on the average score, sorted by Δ=Self−Other. Confounds are not controlled; per-dimension results are in Appendix F.
JudgeSelfOtherΔ
Kimi K2.64.883.79+1.09
GPT-5.55.784.92+0.86
DeepSeek-V4-Pro4.864.34+0.52
Claude Opus 4.74.013.50+0.51
Qwen3.6-Plus4.734.49+0.24
Gemini 3.1 Pro4.414.21+0.21
Llama 4 Maverick5.515.47+0.04
Mistral Large 35.946.02−0.08
Grok 4.33.684.35−0.67
Qwen3.6-35B-A3B3.754.65−0.90
Figure 3: Selection patterns by model (300 selections), projected via Multidimensional Scaling (Borg and Groenen, 2005). Each subplot highlights one model’s selections against the full set; models are ordered by within-model Jaccard similarity (descending).
Figure 3: Selection patterns by model (300 selections), projected via Multidimensional Scaling (Borg and Groenen, 2005). Each subplot highlights one model’s selections against the full set; models are ordered by within-model Jaccard similarity (descending).
Table 3: Confound-controlled self-preference coefficient βS from the mixed-effects model, before and after excluding the two low-discrimination judges. Avg. = the average of the four rubric dimensions and the primary outcome; Org. = Originality, Dim. = Dimensionality, Coh. = Coherence, Tel. = Tellability. The positive effects on the average score and three dimensions vanish once low-discrimination judges are removed; only a negative Originality effect persists, which runs counter to self-preference.
Full (10 judges)Excl. low-discrim. (8)
Dim.βS95% CIβS95% CI
Avg.+0.181∗[ 0.13, 0.23 ]+0.003[−0.05, 0.06 ]
Org.−0.144∗[−0.21, −0.08]−0.133∗[−0.21, −0.05]
Dim.+0.316∗[ 0.24, 0.39 ]+0.039[−0.05, 0.12 ]
Coh.+0.276∗[ 0.21, 0.34 ]+0.042[−0.03, 0.12 ]
Tel.+0.278∗[ 0.21, 0.34 ]+0.065[−0.01, 0.14 ]
Figure 4: Score frequency distribution across all Experiment 1 (blind) evaluations, used as a sanity check. Each cell reports the percentage of evaluations in which a judge assigned that score for the given dimension.
Figure 4: Score frequency distribution across all Experiment 1 (blind) evaluations, used as a sanity check. Each cell reports the percentage of evaluations in which a judge assigned that score for the given dimension.
Table 4: Per-condition descriptive statistics for Experiment 2. Mean (SD) across all judges, pairs, and repetitions. TL = label matches actual source; FL = label contradicts actual source.
ActualConditionAverageOrg.Dim.Coh.Tell.
SelfTL-self4.92 (1.04)4.76 (1.24)5.16 (1.55)3.85 (1.74)5.91 (1.09)
-producedFL-other4.50 (1.10)4.28 (1.25)4.60 (1.60)3.56 (1.69)5.55 (1.14)
OtherFL-self4.91 (1.06)4.83 (1.23)5.09 (1.57)3.92 (1.80)5.79 (1.12)
-producedTL-other4.48 (1.16)4.30 (1.19)4.52 (1.68)3.63 (1.82)5.47 (1.22)
Figure 5: Per-judge mean deviation from the Experiment 1 (no-label) baseline under self-labels (red) and other-labels (blue) on the average score, with 95 percent CIs.
Figure 5: Per-judge mean deviation from the Experiment 1 (no-label) baseline under self-labels (red) and other-labels (blue) on the average score, with 95 percent CIs.
Table 5: Fixed-effect estimates from the mixed-effects models for Experiment 2. βL captures the displayed-label effect; βA the actual-source effect (expected ≈0 under successful quality matching); βL​A their interaction. p∗⁣∗∗<.001. Across all dimensions, βL is the only significant fixed effect, and this pattern is robust to excluding two low-discrimination judges (see Appendix H).
Dim.EffectEstimate95% CI𝐩
βL+0.43∗⁣∗∗[+0.39,+0.47]<.001
Avg.βA+0.02[−0.09,+0.14].689
βL​A−0.01[−0.06,+0.05].823
βL+0.53∗⁣∗∗[+0.47,+0.60]<.001
Org.βA−0.03[−0.15,+0.09].663
βL​A−0.05[−0.14,+0.05].346
βL+0.57∗⁣∗∗[+0.50,+0.65]<.001
Dim.βA+0.11[−0.06,+0.27].197
βL​A−0.02[−0.12,+0.08].721
βL+0.29∗⁣∗∗[+0.24,+0.35]<.001
Coh.βA−0.06[−0.26,+0.14].541
βL​A−0.01[−0.09,+0.07].805
βL+0.32∗⁣∗∗[+0.26,+0.38]<.001
Tel.βA+0.07[−0.05,+0.20].257
βL​A+0.05[−0.04,+0.13].261
Figure 9: Average score assigned by each judge (rows) to each selector’s passages (columns). Diagonal cells indicate self-evaluation. Row-level variation reflects judge-level severity differences; column-level variation reflects selection quality. Selectors that receive consistently high scores across judges—notably GPT-5.5 and Kimi K2.6—also score highly on the diagonal, suggesting that the raw self–other gap is driven in part by selection quality rather than genuine self-preference.
Figure 9: Average score assigned by each judge (rows) to each selector’s passages (columns). Diagonal cells indicate self-evaluation. Row-level variation reflects judge-level severity differences; column-level variation reflects selection quality. Selectors that receive consistently high scores across judges—notably GPT-5.5 and Kimi K2.6—also score highly on the diagonal, suggesting that the raw self–other gap is driven in part by selection quality rather than genuine self-preference.
Table 6: Per-judge mean deviation from the Experiment 1 (no-label) baseline under self- and other-label conditions on the average score. Patterns: symmetric (self ↑, other ↓), self-boost (self ↑ only), other-penalty (other ↓ only), general-penalty (both ↓). Rows grouped by pattern; within each group, judges are sorted by the magnitude of the label-induced shift. p∗<.05 (one-sample t-test, H0: μ=0).
JudgeSelf-labelOther-labelPattern
Qwen3.6-35B-A3B+0.131∗−0.523∗symmetric
Gemini 3.1 Pro+0.302∗−0.339∗symmetric
Qwen3.6-Plus+0.175∗−0.444∗symmetric
Kimi K2.6+0.147∗−0.386∗symmetric
Mistral Large 3+0.081∗−0.265∗symmetric
Grok 4.3+0.126∗−0.195∗symmetric
DeepSeek-V4-Pro+0.443∗+0.101self-boost
Claude Opus 4.7+0.275∗−0.023self-boost
GPT-5.5−0.019−0.340∗other-penalty
Llama 4 Maverick−0.192∗−0.384∗general-penalty
Table 7: All models were accessed via OpenRouter222https://openrouter.ai; exact snapshot identifiers are provided. The set spans nine developers and covers frontier general-purpose systems alongside open-weight models of varying scale and training lineage, providing sufficient heterogeneity to probe self-preference across diverse judges and model families. All models are queried with temperature = 1.0 and reasoning_effort = high where each parameter is supported; unsupported parameters were left at their defaults.
ModelOpenRouter ID
Commercial
Claude Opus 4.7anthropic/claude-4.7-opus-20260416
Gemini 3.1 Progoogle/gemini-3.1-pro-preview-20260219
GPT-5.5openai/gpt-5.5-20260423
Grok 4.3x-ai/grok-4.3-20260430
Qwen3.6-Plusqwen/qwen3.6-plus-04-02
Open-weight
DeepSeek-V4-Prodeepseek/deepseek-v4-pro-20260423
Kimi K2.6moonshotai/kimi-k2.6-20260420
Llama 4 Maverickmeta-llama/llama-4-maverick-17b-128e-instruct
Mistral Large 3mistralai/mistral-large-2512
Qwen3.6-35B-A3Bqwen/qwen3.6-35b-a3b-20260415
Table 8: Four-dimension evaluation rubric with anchor descriptions at the endpoints of a 7-point Likert scale. Theoretical grounding for each dimension is provided in the text below.
DimensionAnchor at 1 (Low)Anchor at 7 (High)
OriginalityThe selected constraints lean on familiar story patterns and well-worn tropes, producing a combination that is predictable and conventional.The selected constraints break from familiar story patterns and well-worn tropes, producing a combination that is unpredictable and fresh.
DimensionalityThe selected constraints operate independently; altering one does not affect the meaning or function of the others.The selected constraints are mutually constitutive, each element defining the others’ meaning and weight, producing narrative depth.
CoherenceThe selected constraints contain irreconcilable contradictions that no narrative device can resolve within a single coherent story.The selected constraints integrate seamlessly into a single narrative without logical tensions or unresolved contradictions.
TellabilityThe selected constraints lack a point worth telling—no conflict, question, or stake arises that would make an audience want to hear or tell this story.The selected constraints carry a point worth telling—a clear conflict, urgent question, or meaningful stake that naturally invites narration and sustains attention.
Table 9: Leave-one-out k-NN classification of selection profiles by source model (N=300, chance =0.10, 10 classes). Cohen’s κ ranges from 0.41 to 0.45, indicating moderate chance-corrected agreement. k=5 is used as the primary configuration. p∗⁣∗∗<.001 (binomial test against H0:p=1/10).
kAccuracyCorrect/NMacro-F1κp
10.477143/3000.4270.419***
30.467140/3000.4050.407***
50.507152/3000.4480.452***
70.493148/3000.4240.437***
Table 10: Per-model k-NN classification report (k=5, n=30 per model), sorted by F1 descending. Two asymmetric patterns stand out: Llama 4 Maverick and GPT-5.5 achieve perfect recall but lower precision, whereas Mistral Large 3 achieves perfect precision but is rarely selected as the prediction.
ModelTPFPFNPrec.Rec.F1
Claude Opus 4.7251050.710.830.77
Llama 4 Maverick302200.581.000.73
GPT-5.5303700.451.000.62
Kimi K2.62018100.530.670.59
Qwen3.6-Plus2024100.450.670.54
DeepSeek-V4-Pro76230.540.230.33
Gemini 3.1 Pro812220.400.270.32
Qwen3.6-35B-A3B510250.330.170.22
Grok 4.349260.310.130.19
Mistral Large 330271.000.100.18
Table 11: Within-model selection consistency measured by pairwise Jaccard similarity, computed over all (302)=435 unique pairs from 30 selection runs per model. The random null is the expected mean Jaccard under uniform random selection from 200 constraints (10,000 permutations; mean =0.054, 95% CI: [0.051,0.058]). ∗ denotes p<.0001 relative to the random null (one-tailed permutation test).
ModelMeanSDMinMax
Commercial
Claude Opus 4.70.2180.0860.0260.482
Gemini 3.1 Pro0.1900.0930.0260.600
GPT-5.50.4050.0920.1430.667
Grok 4.30.1200.0600.0000.333
Qwen3.6-Plus0.2350.0790.0260.539
Open-weight
DeepSeek-V4-Pro0.1360.0620.0000.429
Kimi K2.60.2210.0690.0810.429
Llama 4 Maverick0.2470.0910.0530.600
Mistral Large 30.0720.0400.0000.212
Qwen3.6-35B-A3B0.1530.0640.0000.379
Random null0.054
Table 12: Raw self–other comparison per judge across all four rubric dimensions. Δ=Self−Other. Confounds are not controlled.
OriginalityDimensionalityCoherenceTellability
JudgeSelfOtherΔSelfOtherΔSelfOtherΔSelfOtherΔ
Kimi K2.63.434.47−1.045.443.40+2.054.722.44+2.285.934.85+1.08
GPT-5.54.464.74−0.285.995.16+0.835.833.89+1.946.825.87+0.95
DeepSeek-V4-Pro5.685.07+0.604.404.02+0.384.023.48+0.545.334.78+0.56
Claude Opus 4.73.843.87−0.034.133.23+0.912.692.13+0.565.364.77+0.58
Qwen3.6-Plus4.745.00−0.264.964.48+0.473.533.16+0.375.695.30+0.39
Gemini 3.1 Pro3.764.63−0.884.944.34+0.602.862.27+0.596.105.59+0.51
Llama 4 Maverick5.394.89+0.505.895.83+0.064.905.37−0.475.865.78+0.07
Mistral Large 35.985.70+0.286.266.38−0.125.005.41−0.416.526.60−0.07
Grok 4.34.974.42+0.553.024.07−1.042.063.51−1.454.695.41−0.72
Qwen3.6-35B-A3B3.734.62−0.893.964.93−0.972.293.47−1.185.025.59−0.56
Table 13: Robustness check: fixed-effect estimates from the mixed-effects models for Experiment 2 after excluding the two low-discrimination judges (Llama 4 Maverick and Mistral Large 3). Model specification identical to Table 5. βL captures the displayed-label effect; βA the actual-source effect (expected ≈0 under successful quality matching); βL​A their interaction. p∗⁣∗∗<.001. The label effect (βL) remains the only significant fixed effect on every dimension, and its magnitude is comparable to or slightly larger than the full-sample estimate, confirming that the main findings are not driven by the two excluded judges.
Dim.EffectEstimate95% CI𝐩
βL+0.47∗⁣∗∗[+0.43,+0.52]<.001
Avg.βA+0.08[−0.07,+0.22].293
βL​A−0.02[−0.09,+0.05].638
βL+0.57∗⁣∗∗[+0.49,+0.65]<.001
Org.βA−0.12[−0.27,+0.02].089
βL​A−0.03[−0.14,+0.08].595
βL+0.68∗⁣∗∗[+0.59,+0.76]<.001
Dim.βA+0.15[−0.06,+0.35].164
βL​A−0.04[−0.17,+0.08].468
βL+0.27∗⁣∗∗[+0.20,+0.33]<.001
Coh.βA+0.15[−0.09,+0.39].214
βL​A−0.03[−0.12,+0.06].538
βL+0.38∗⁣∗∗[+0.31,+0.45]<.001
Tel.βA+0.13[−0.03,+0.28].101
βL​A+0.04[−0.06,+0.13].462

왜 중요한가

AI가 다른 AI의 답을 자동으로 채점하는 방식(LLM-as-a-judge)이 널리 쓰이는 상황에서, 실제 품질이 아니라 '이게 누구 것이다'라는 표시만으로 점수가 바뀔 수 있다는 것은 평가 시스템의 신뢰성에 직접적인 위협이 된다. 특히 문체나 실제 품질 같은 다른 요인을 다 제거하고도 순수하게 라벨 하나로 점수가 오르내렸다는 점에서, AI 평가 결과를 그대로 신뢰하기 전에 저자 정보 노출 여부를 반드시 점검해야 한다는 실무적 시사점을 준다.

이 논문의 용어

  • LLM-as-a-judge · 사람 대신 대형언어모델이 다른 AI의 답변이나 결과물을 채점·평가하는 방식
  • 자기선호(self-preference) · AI 채점자가 자기 자신이 만든 결과물을 더 좋게 평가하는 편향 현상
  • 블라인드 평가 · 누가 만든 결과물인지 알려주지 않고 채점하는 조건
  • 혼합효과모델(mixed-effects model) · 채점자 성향, 선택 품질 등 여러 변수가 뒤섞인 영향을 통계적으로 분리해내는 분석 기법
  • 저자 표시 라벨(authorship label) · 결과물의 실제 제작자가 아니라 채점 시점에 '이것은 누가 만든 것'이라고 알려주는 표시

저자 · Songeun Chae, Min Kim, Donghoon Jung, Seojin Choi, Seohyon Jung

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL MEDIA 최신 기사

그림 출처: Songeun Chae et al., arXiv:2608.18091, CC BY 4.0