K-文化的一切——从回归到 K-美妆,发送到您的邮箱订阅邮件

METAL MEDIA

Institutional Prestige as Geographic Bias in Large Language Models: Evidence from Three Factorial Experiments with Bootstrap Confidence Intervals

arXiv:2608.181072026-08-20

AI给候选人打分时,姓名影响不大,但学校名气和论文发表刊物却能左右分数

研究者让四款大型语言模型扮演评审,对奖学金、招聘、贷款、科研经费和政策提案等五类场景中的虚拟候选人打分,发现姓名所暗示的族裔背景几乎不影响分数,但毕业院校的声望和论文发表期刊的声望却会显著左右评分。在Nature上发表一篇论文带来的加分,是就读顶尖大学带来加分的5.7倍。研究者还提出一个新指标NBI,证明低声望背景的候选人不仅平均分更低,评分本身也更不稳定。

METAL MEDIA 解读图

AI给候选人打分时,姓名影响不大,但学校名气和论文发表刊物却能左右分数

  1. 01研究对Claude Haiku 4.5、GPT-4o-mini、Gemini 2.0 Flash和Llama 3.1 8B四款模型共进行了4,320次API调用,在奖学金、招聘、贷款审批、科研经费、公共政策提案五个领域中,系统改变候选人姓名(英语系/拉丁裔/阿拉伯裔)、就读院校等级、所在国家、发表期刊等变量
  2. 02在10分制评分中,姓名族裔差异在统计上不显著(±0.094,置信区间跨越零),但院校声望从顶级(MIT级)降到无排名院校,平均分下降0.297分,该效应在统计上稳健
  3. 03将院校声望(+0.185)与来源国(发达/发展中,+0.126)的影响分开检验后发现,声望效应是国家效应的1.5倍;声望高的拉美名校(UNAM)并未被评得比一所无排名的美国院校(Framingham State)低,排除了单纯的国家偏见
  4. 04将期刊声望(Nature对比一个边缘开放获取期刊)与院校声望交叉实验后发现,期刊效应(+1.937)是院校效应(+0.341)的5.7倍;在Nature发表论文对来自低声望院校(瓜亚基尔大学)候选人的加分,甚至超过对MIT候选人的加分,形成一种'补偿效应'
  5. 05研究者提出的中智偏见指数(NBI)显示,低声望背景候选人面临双重不利:不仅平均得分更低,在相同条件下评分的波动(不确定性)也更大,这是只看平均分的常规审计方法无法发现的隐性劣势
这是 METAL MEDIA 制作的解读图,并非论文作者提供的原图。

他们做了什么

  1. 研究对Claude Haiku 4.5、GPT-4o-mini、Gemini 2.0 Flash和Llama 3.1 8B四款模型共进行了4,320次API调用,在奖学金、招聘、贷款审批、科研经费、公共政策提案五个领域中,系统改变候选人姓名(英语系/拉丁裔/阿拉伯裔)、就读院校等级、所在国家、发表期刊等变量
  2. 在10分制评分中,姓名族裔差异在统计上不显著(±0.094,置信区间跨越零),但院校声望从顶级(MIT级)降到无排名院校,平均分下降0.297分,该效应在统计上稳健
  3. 将院校声望(+0.185)与来源国(发达/发展中,+0.126)的影响分开检验后发现,声望效应是国家效应的1.5倍;声望高的拉美名校(UNAM)并未被评得比一所无排名的美国院校(Framingham State)低,排除了单纯的国家偏见
  4. 将期刊声望(Nature对比一个边缘开放获取期刊)与院校声望交叉实验后发现,期刊效应(+1.937)是院校效应(+0.341)的5.7倍;在Nature发表论文对来自低声望院校(瓜亚基尔大学)候选人的加分,甚至超过对MIT候选人的加分,形成一种'补偿效应'
  5. 研究者提出的中智偏见指数(NBI)显示,低声望背景候选人面临双重不利:不仅平均得分更低,在相同条件下评分的波动(不确定性)也更大,这是只看平均分的常规审计方法无法发现的隐性劣势
Figure 1: Study 1 – Institution-Tier Gradient by Model with 95% Bootstrap CI (cross-model).
Figure 1: Study 1 – Institution-Tier Gradient by Model with 95% Bootstrap CI (cross-model).
Table 1: Study 1 – Mean Score by Institution Tier. Cross-model CI from 10,000 bootstrap iterations. ✓ = 95% CI entirely positive.
ModelT1 MITT2 UChileT3 UNALT5 UGyeGradient T1−T595% CI
Claude Haiku 4.57.4337.2337.2227.133+0.300[+0.111,+0.478]
GPT-4o-mini8.3788.2338.2118.156+0.222[+0.022,+0.422]
Gemini 2.0 Flash7.5567.3787.3117.189+0.367[+0.133,+0.611]
Llama 3.1 8B7.6787.4447.589∗7.378+0.300[+0.044,+0.556]
Cross-model7.7617.5727.5837.464+0.297[+0.175,+0.422] ✓
Figure 2: Study 2 – 2×2 Prestige × Country Cell Means by Model.
Figure 2: Study 2 – 2×2 Prestige × Country Cell Means by Model.
Table 2: Study 1 – Mean Score by Name Origin. Bootstrap CIs cross zero for all contrasts (⇒ non-significant).
ModelAngloLatinoArabicMax gapSignificant?
Claude Haiku 4.57.2087.2337.3250.117No
GPT-4o-mini8.2178.2508.2670.050No
Gemini 2.0 Flash7.2677.4177.3920.150No
Llama 3.1 8B7.4677.5507.5500.083No
Cross-model7.5407.6127.6330.094No (CI ∋ 0)
Figure 3: Study 3 – 2×2 Journal × Institution Prestige. The “rescue effect” cell (UGye + Nature) shows the largest journal premium (Δ=+2.13), indicating that Nature publication compensates for low institutional prestige more than for high institutional prestige.
Figure 3: Study 3 – 2×2 Journal × Institution Prestige. The “rescue effect” cell (UGye + Nature) shows the largest journal premium (Δ=+2.13), indicating that Nature publication compensates for low institutional prestige more than for high institutional prestige.
Table 3: Study 2 – 2×2 Cell Means and Factorial Effects with 95% Bootstrap CIs. † = CI entirely positive (significant). Dev = Developed; Dvlp = Developing.
ModelMITUNAMFSUUGyePrestige (95% CI)Country (95% CI)
(Hi,Dev)(Hi,Dvlp)(Lo,Dev)(Lo,Dvlp)
Haiku 4.57.4117.2337.0117.144+0.244† [+0.106,+0.383]+0.022 [−0.117,+0.167]
GPT-4o-mini8.3898.2228.1788.156+0.139† [+0.000,+0.278]+0.094 [−0.044,+0.233]
Gemini 2.0F7.5787.3007.3007.167+0.206† [+0.039,+0.372]+0.206† [+0.044,+0.372]
Llama 3.1 8B7.6567.3447.3787.322+0.150 [−0.045,+0.344]+0.183 [−0.017,+0.378]
Cross-model7.7587.5257.4677.447+0.185† [+0.093,+0.275]+0.126† [+0.037,+0.218]
Table 4: Study 2 – Critical Contrast: UNAM (Mexico, high prestige) vs. Framingham State (USA, low prestige).
ModelUNAMFramingham St.UNAM−FSUInterpretation
Claude Haiku 4.57.2337.011+0.222Prestige wins
GPT-4o-mini8.2228.178+0.044≈ Equal
Gemini 2.0 Flash7.3007.300±0.000≈ Equal
Llama 3.1 8B7.3447.378−0.033≈ Equal
Cross-model7.5257.467+0.058[−0.072,+0.186] – trend, n.s.
Table 5: Study 2 – Prestige vs. Country Effect by Domain with 95% Bootstrap CIs. † = CI entirely positive. ∗⁣∗ = normatively unjustified (no legitimate weight under anti-discrimination principles).
DomainPrestige effect (95% CI)Country effect (95% CI)DominantNote
Scholarship+0.160​[−0.021,+0.340]+0.090​[−0.090,+0.271]Partially justified
Hiring+0.160†​[+0.035,+0.292]+0.187†​[+0.062,+0.312]CountryPartially justified
Credit+0.278†​[+0.076,+0.479]+0.264†​[+0.062,+0.465]PrestigeNot justified ∗⁣∗
Health+0.125​[−0.076,+0.326]+0.097​[−0.104,+0.299]Partially justified
Public Policy+0.201†​[+0.028,+0.375]−0.007​[−0.181,+0.167]PrestigeNot justified ∗⁣∗
Table 6: Study 3 – 2×2 Journal × Institution Prestige Cell Means. All four models, 5 domains, 3 names, 6 reps per cell. † = 95% bootstrap CI entirely positive.
Nature (hi-journal)NCML (lo-journal)
MIT (hi-inst)UGye (lo-inst)MIT (hi-inst)UGye (lo-inst)
Mean score7.9717.8226.2255.694
Journal effect (institution row)ΔMIT=+1.746†​[+1.564,+1.925]ΔUGye=+2.128†​[+1.950,+2.297]
Cross-cell main effects:
Journal (Nature−NCML)+1.937† [+1.811,+2.062]
Institution (MIT−UGye)+0.341† [+0.184,+0.504]
Ratio journal/institution5.7×
Interaction (rescue effect)−0.382 (Nature rescues UGye more than MIT)
Table 7: Study 3 – Per-Model Journal and Institution Prestige Effects with 95% CIs. † = CI entirely positive.
ModelJournal effect (95% CI)Institution effect (95% CI)
Claude Haiku 4.5+2.937†​[+2.733,+3.139]+0.232​[−0.128,+0.594]
GPT-4o-mini+1.322†​[+1.139,+1.506]+0.099​[−0.133,+0.328]
Gemini 2.0 Flash+2.611†​[+2.372,+2.844]+0.598†​[+0.244,+0.956]
Llama 3.1 8B+0.875†​[+0.694,+1.053]+0.430†​[+0.231,+0.631]
Cross-model+1.937†​[+1.811,+2.062]+0.341†​[+0.184,+0.504]
Table 8: NBI ⟨T,I,F⟩ for Reference (Anglo-MIT) and T5 Profiles.
ModelProfileTIFInterpretation
Haiku 4.5Anglo T1 (ref)0.7400.0790.000Reference
Anglo T50.7100.1180.030Prestige penalty
Latino T50.7070.1090.033Prestige penalty
Arabic T50.7230.1090.017Prestige penalty
GPT-4o-miniAnglo T1 (ref)0.8280.1260.000Reference
Anglo T50.8110.1260.017Prestige penalty
Latino T50.8090.1250.020Prestige penalty
Arabic T50.8110.1260.017Prestige penalty
Gemini 2.0FAnglo T1 (ref)0.7530.1390.000Reference
Anglo T50.7050.1260.046Prestige penalty
Latino T50.7230.1080.033Prestige penalty
Arabic T50.7270.1040.027Prestige penalty
Llama 3.1 8BAnglo T1 (ref)0.7800.1150.000Reference
Anglo T50.7220.1870.057Strongest bias
Latino T50.7470.1530.033Prestige penalty
Arabic T50.7430.1340.037Prestige penalty

为什么重要

随着AI评审工具越来越多地被用于奖学金评审、招聘、信贷审批和科研经费分配等场景,如果只检测姓名带来的偏见就判定系统'公平',可能会忽略院校声望和发表刊物带来的更强偏见。这项研究用实证方式揭示了非名校或发展中国家研究者可能面临的结构性不利。

本文术语

  • 自助法置信区间(bootstrap confidence interval) · 通过反复对数据重新抽样来估计某个效应是否可靠的统计方法
  • 中智偏见指数(Neutrosophic Bias Index, NBI) · 本文提出的新指标,用真值、不确定性、假值三个数值同时刻画平均分差异和评分不稳定性
  • 因子设计(factorial design) · 同时改变姓名、院校等级、国家、期刊等多个变量的实验设计方法,用于分离出每个变量各自的独立影响
  • 补偿效应(rescue effect) · 指在顶级期刊发表论文能够弥补院校声望不足的现象,且这种弥补对低声望院校候选人的效果比对高声望院校候选人更强

论文原文摘要(英文)

We investigate whether large language models (LLMs) systematically discriminate in candidate evaluations based on applicant name ethnicity and/or institutional prestige and geographic location. Three factorial experiments are reported (4,320 API calls, four LLMs, five professional domains). Study 1 (3x4 design) finds a statistically robust institution-tier gradient of +0.297 points on a 10-point scale (95% bootstrap CI: +0.175 to +0.422), while name-origin effects are negligible and non-significant (95% CI crosses zero). Study 2 (2x2 Prestige x Country design) breaks the prestige-geography confound: the prestige effect (+0.185; 95% CI: +0.093 to +0.275) exceeds the country-of-origin effect (+0.126; 95% CI: +0.037 to +0.218) by 1.5x. Study 3 (2x2 Journal x Institution design) reveals that journal prestige (Nature vs. a peripheral open-access journal) dominates institutional prestige by 5.7x: journal effect +1.937 (95% CI: +1.811 to +2.062) vs. institution effect +0.341 (95% CI: +0.184 to +0.504). A "rescue effect" is confirmed: publishing in Nature compensates for low institutional prestige more strongly for candidates from the University of Guayaquil (+2.127) than from MIT (+1.745). Results are quantified using the Neutrosophic Bias Index NBI ; the I component reveals elevated evaluation inconsistency for low-prestige profiles, an epistemic disadvantage not captured by mean-only metrics. Code and data: https://github.com/mleyvaz/geo-bias-llm

作者 · Maikel Leyva-Vazquez, Florentin Smarandache

在 arXiv 阅读

最新论文

全部论文 →

METAL MEDIA 最新报道

图片来源: Maikel Leyva-Vazquez et al., arXiv:2608.18107, CC BY 4.0