Everything K-culture — comebacks to K-beauty, straight to your inboxGet it in your inbox›
Institutional Prestige as Geographic Bias in Large Language Models: Evidence from Three Factorial Experiments with Bootstrap Confidence Intervals
arXiv:2608.181072026-08-20
In AI candidate scoring, school prestige and where you publish matter far more than your name
Testing four large language models as automated evaluators across scholarship, hiring, credit, health and policy scenarios, researchers found that the ethnic signal carried by a candidate's name barely moved scores, but institutional prestige and journal prestige did, strongly. A single publication in Nature swung scores 5.7 times more than attending a top-tier university did. The team introduced a new Neutrosophic Bias Index (NBI) showing that low-prestige candidates are hurt twice over: lower average scores and more inconsistent evaluations.
METAL MEDIA explanatory visual
In AI candidate scoring, school prestige and where you publish matter far more than your name
01Across 4,320 API calls to Claude Haiku 4.5, GPT-4o-mini, Gemini 2.0 Flash, and Llama 3.1 8B, the study ran three factorial experiments varying applicant name origin (Anglo/Latino/Arabic), institution prestige tier, country, and journal prestige across five professional domains
02Name-origin effects on 10-point scores were statistically non-significant (±0.094, confidence interval crosses zero), but institution tier produced a robust gradient of +0.297 points from top-ranked (MIT-level) to unranked institutions
03Separating institutional prestige (+0.185) from country of origin (developed vs developing, +0.126) showed prestige mattered 1.5 times more than country, and a prestigious Latin American university (UNAM) was not rated below an unranked US university (Framingham State), ruling out pure country bias
04Crossing journal prestige (Nature vs. a peripheral open-access journal) with institutional prestige revealed the journal effect (+1.937) was 5.7 times larger than the institution effect (+0.341); publishing in Nature boosted scores for candidates from a low-prestige university (Universidad de Guayaquil) even more than for MIT candidates, a 'rescue effect'
05The newly introduced Neutrosophic Bias Index (NBI) showed low-prestige candidates face a compound disadvantage: not just lower average scores but also higher evaluation inconsistency, a form of epistemic disadvantage that mean-score-only audits miss
An explanatory diagram made by METAL MEDIA, not a figure supplied by the paper's authors.
What they did
Across 4,320 API calls to Claude Haiku 4.5, GPT-4o-mini, Gemini 2.0 Flash, and Llama 3.1 8B, the study ran three factorial experiments varying applicant name origin (Anglo/Latino/Arabic), institution prestige tier, country, and journal prestige across five professional domains
Name-origin effects on 10-point scores were statistically non-significant (±0.094, confidence interval crosses zero), but institution tier produced a robust gradient of +0.297 points from top-ranked (MIT-level) to unranked institutions
Separating institutional prestige (+0.185) from country of origin (developed vs developing, +0.126) showed prestige mattered 1.5 times more than country, and a prestigious Latin American university (UNAM) was not rated below an unranked US university (Framingham State), ruling out pure country bias
Crossing journal prestige (Nature vs. a peripheral open-access journal) with institutional prestige revealed the journal effect (+1.937) was 5.7 times larger than the institution effect (+0.341); publishing in Nature boosted scores for candidates from a low-prestige university (Universidad de Guayaquil) even more than for MIT candidates, a 'rescue effect'
The newly introduced Neutrosophic Bias Index (NBI) showed low-prestige candidates face a compound disadvantage: not just lower average scores but also higher evaluation inconsistency, a form of epistemic disadvantage that mean-score-only audits miss
Figure 1: Study 1 – Institution-Tier Gradient by Model with 95% Bootstrap CI (cross-model).
Table 1: Study 1 – Mean Score by Institution Tier. Cross-model CI from 10,000 bootstrap iterations. ✓ = 95% CI entirely positive.
Model
T1 MIT
T2 UChile
T3 UNAL
T5 UGye
Gradient T1−T5
95% CI
Claude Haiku 4.5
7.433
7.233
7.222
7.133
+0.300
[+0.111,+0.478]
GPT-4o-mini
8.378
8.233
8.211
8.156
+0.222
[+0.022,+0.422]
Gemini 2.0 Flash
7.556
7.378
7.311
7.189
+0.367
[+0.133,+0.611]
Llama 3.1 8B
7.678
7.444
7.589∗
7.378
+0.300
[+0.044,+0.556]
Cross-model
7.761
7.572
7.583
7.464
+0.297
[+0.175,+0.422] ✓
Figure 2: Study 2 – 2×2 Prestige × Country Cell Means by Model.
Table 2: Study 1 – Mean Score by Name Origin. Bootstrap CIs cross zero for all contrasts (⇒ non-significant).
Model
Anglo
Latino
Arabic
Max gap
Significant?
Claude Haiku 4.5
7.208
7.233
7.325
0.117
No
GPT-4o-mini
8.217
8.250
8.267
0.050
No
Gemini 2.0 Flash
7.267
7.417
7.392
0.150
No
Llama 3.1 8B
7.467
7.550
7.550
0.083
No
Cross-model
7.540
7.612
7.633
0.094
No (CI ∋ 0)
Figure 3: Study 3 – 2×2 Journal × Institution Prestige. The “rescue effect” cell (UGye + Nature) shows the largest journal premium (Δ=+2.13), indicating that Nature publication compensates for low institutional prestige more than for high institutional prestige.
Table 3: Study 2 – 2×2 Cell Means and Factorial Effects with 95% Bootstrap CIs. † = CI entirely positive (significant). Dev = Developed; Dvlp = Developing.
Model
MIT
UNAM
FSU
UGye
Prestige (95% CI)
Country (95% CI)
(Hi,Dev)
(Hi,Dvlp)
(Lo,Dev)
(Lo,Dvlp)
Haiku 4.5
7.411
7.233
7.011
7.144
+0.244† [+0.106,+0.383]
+0.022 [−0.117,+0.167]
GPT-4o-mini
8.389
8.222
8.178
8.156
+0.139† [+0.000,+0.278]
+0.094 [−0.044,+0.233]
Gemini 2.0F
7.578
7.300
7.300
7.167
+0.206† [+0.039,+0.372]
+0.206† [+0.044,+0.372]
Llama 3.1 8B
7.656
7.344
7.378
7.322
+0.150 [−0.045,+0.344]
+0.183 [−0.017,+0.378]
Cross-model
7.758
7.525
7.467
7.447
+0.185† [+0.093,+0.275]
+0.126† [+0.037,+0.218]
Table 4: Study 2 – Critical Contrast: UNAM (Mexico, high prestige) vs. Framingham State (USA, low prestige).
Model
UNAM
Framingham St.
UNAM−FSU
Interpretation
Claude Haiku 4.5
7.233
7.011
+0.222
Prestige wins
GPT-4o-mini
8.222
8.178
+0.044
≈ Equal
Gemini 2.0 Flash
7.300
7.300
±0.000
≈ Equal
Llama 3.1 8B
7.344
7.378
−0.033
≈ Equal
Cross-model
7.525
7.467
+0.058
[−0.072,+0.186] – trend, n.s.
Table 5: Study 2 – Prestige vs. Country Effect by Domain with 95% Bootstrap CIs. † = CI entirely positive. ∗∗ = normatively unjustified (no legitimate weight under anti-discrimination principles).
Domain
Prestige effect (95% CI)
Country effect (95% CI)
Dominant
Note
Scholarship
+0.160[−0.021,+0.340]
+0.090[−0.090,+0.271]
—
Partially justified
Hiring
+0.160†[+0.035,+0.292]
+0.187†[+0.062,+0.312]
Country
Partially justified
Credit
+0.278†[+0.076,+0.479]
+0.264†[+0.062,+0.465]
Prestige
Not justified ∗∗
Health
+0.125[−0.076,+0.326]
+0.097[−0.104,+0.299]
—
Partially justified
Public Policy
+0.201†[+0.028,+0.375]
−0.007[−0.181,+0.167]
Prestige
Not justified ∗∗
Table 6: Study 3 – 2×2 Journal × Institution Prestige Cell Means. All four models, 5 domains, 3 names, 6 reps per cell. † = 95% bootstrap CI entirely positive.
Nature (hi-journal)
NCML (lo-journal)
MIT (hi-inst)
UGye (lo-inst)
MIT (hi-inst)
UGye (lo-inst)
Mean score
7.971
7.822
6.225
5.694
Journal effect (institution row)
ΔMIT=+1.746†[+1.564,+1.925]
ΔUGye=+2.128†[+1.950,+2.297]
Cross-cell main effects:
Journal (Nature−NCML)
+1.937† [+1.811,+2.062]
Institution (MIT−UGye)
+0.341† [+0.184,+0.504]
Ratio journal/institution
5.7×
Interaction (rescue effect)
−0.382 (Nature rescues UGye more than MIT)
Table 7: Study 3 – Per-Model Journal and Institution Prestige Effects with 95% CIs. † = CI entirely positive.
Model
Journal effect (95% CI)
Institution effect (95% CI)
Claude Haiku 4.5
+2.937†[+2.733,+3.139]
+0.232[−0.128,+0.594]
GPT-4o-mini
+1.322†[+1.139,+1.506]
+0.099[−0.133,+0.328]
Gemini 2.0 Flash
+2.611†[+2.372,+2.844]
+0.598†[+0.244,+0.956]
Llama 3.1 8B
+0.875†[+0.694,+1.053]
+0.430†[+0.231,+0.631]
Cross-model
+1.937†[+1.811,+2.062]
+0.341†[+0.184,+0.504]
Table 8: NBI ⟨T,I,F⟩ for Reference (Anglo-MIT) and T5 Profiles.
Model
Profile
T
I
F
Interpretation
Haiku 4.5
Anglo T1 (ref)
0.740
0.079
0.000
Reference
Anglo T5
0.710
0.118
0.030
Prestige penalty
Latino T5
0.707
0.109
0.033
Prestige penalty
Arabic T5
0.723
0.109
0.017
Prestige penalty
GPT-4o-mini
Anglo T1 (ref)
0.828
0.126
0.000
Reference
Anglo T5
0.811
0.126
0.017
Prestige penalty
Latino T5
0.809
0.125
0.020
Prestige penalty
Arabic T5
0.811
0.126
0.017
Prestige penalty
Gemini 2.0F
Anglo T1 (ref)
0.753
0.139
0.000
Reference
Anglo T5
0.705
0.126
0.046
Prestige penalty
Latino T5
0.723
0.108
0.033
Prestige penalty
Arabic T5
0.727
0.104
0.027
Prestige penalty
Llama 3.1 8B
Anglo T1 (ref)
0.780
0.115
0.000
Reference
Anglo T5
0.722
0.187
0.057
Strongest bias
Latino T5
0.747
0.153
0.033
Prestige penalty
Arabic T5
0.743
0.134
0.037
Prestige penalty
Why it matters
As LLMs get deployed for scholarship, hiring, credit and grant decisions, checking only for name-based bias can give a false sense of fairness while institutional and publication-venue bias go undetected. This work empirically shows how researchers and applicants from lower-prestige institutions or the Global South could be systematically disadvantaged by AI evaluators.
Terms in this paper
bootstrap confidence interval · a statistical range estimated by repeatedly resampling the data to gauge how reliable a measured effect is
Neutrosophic Bias Index (NBI) · a new metric proposed in this paper expressing bias as three values, truth, indeterminacy, and falsity, capturing both average score penalty and evaluation inconsistency
factorial design · an experimental setup that varies multiple factors (like name, institution tier, country, journal) at once to isolate each factor's independent effect
rescue effect · the finding that publishing in a prestigious journal compensates for low institutional prestige, more strongly than for candidates from high-prestige institutions
Original abstract (English)
We investigate whether large language models (LLMs) systematically discriminate in candidate evaluations based on applicant name ethnicity and/or institutional prestige and geographic location. Three factorial experiments are reported (4,320 API calls, four LLMs, five professional domains). Study 1 (3x4 design) finds a statistically robust institution-tier gradient of +0.297 points on a 10-point scale (95% bootstrap CI: +0.175 to +0.422), while name-origin effects are negligible and non-significant (95% CI crosses zero). Study 2 (2x2 Prestige x Country design) breaks the prestige-geography confound: the prestige effect (+0.185; 95% CI: +0.093 to +0.275) exceeds the country-of-origin effect (+0.126; 95% CI: +0.037 to +0.218) by 1.5x. Study 3 (2x2 Journal x Institution design) reveals that journal prestige (Nature vs. a peripheral open-access journal) dominates institutional prestige by 5.7x: journal effect +1.937 (95% CI: +1.811 to +2.062) vs. institution effect +0.341 (95% CI: +0.184 to +0.504). A "rescue effect" is confirmed: publishing in Nature compensates for low institutional prestige more strongly for candidates from the University of Guayaquil (+2.127) than from MIT (+1.745). Results are quantified using the Neutrosophic Bias Index NBI ; the I component reveals elevated evaluation inconsistency for low-prestige profiles, an epistemic disadvantage not captured by mean-only metrics. Code and data: https://github.com/mleyvaz/geo-bias-llm