Figure 1: Study 1 – Institution-Tier Gradient by Model with 95% Bootstrap CI (cross-model).
Table 1: Study 1 – Mean Score by Institution Tier. Cross-model CI from 10,000 bootstrap iterations. ✓ = 95% CI entirely positive.
Model
T1 MIT
T2 UChile
T3 UNAL
T5 UGye
Gradient T1−T5
95% CI
Claude Haiku 4.5
7.433
7.233
7.222
7.133
+0.300
[+0.111,+0.478]
GPT-4o-mini
8.378
8.233
8.211
8.156
+0.222
[+0.022,+0.422]
Gemini 2.0 Flash
7.556
7.378
7.311
7.189
+0.367
[+0.133,+0.611]
Llama 3.1 8B
7.678
7.444
7.589∗
7.378
+0.300
[+0.044,+0.556]
Cross-model
7.761
7.572
7.583
7.464
+0.297
[+0.175,+0.422] ✓
Figure 2: Study 2 – 2×2 Prestige × Country Cell Means by Model.
Table 2: Study 1 – Mean Score by Name Origin. Bootstrap CIs cross zero for all contrasts (⇒ non-significant).
Model
Anglo
Latino
Arabic
Max gap
Significant?
Claude Haiku 4.5
7.208
7.233
7.325
0.117
No
GPT-4o-mini
8.217
8.250
8.267
0.050
No
Gemini 2.0 Flash
7.267
7.417
7.392
0.150
No
Llama 3.1 8B
7.467
7.550
7.550
0.083
No
Cross-model
7.540
7.612
7.633
0.094
No (CI ∋ 0)
Figure 3: Study 3 – 2×2 Journal × Institution Prestige. The “rescue effect” cell (UGye + Nature) shows the largest journal premium (Δ=+2.13), indicating that Nature publication compensates for low institutional prestige more than for high institutional prestige.
Table 3: Study 2 – 2×2 Cell Means and Factorial Effects with 95% Bootstrap CIs. † = CI entirely positive (significant). Dev = Developed; Dvlp = Developing.
Model
MIT
UNAM
FSU
UGye
Prestige (95% CI)
Country (95% CI)
(Hi,Dev)
(Hi,Dvlp)
(Lo,Dev)
(Lo,Dvlp)
Haiku 4.5
7.411
7.233
7.011
7.144
+0.244† [+0.106,+0.383]
+0.022 [−0.117,+0.167]
GPT-4o-mini
8.389
8.222
8.178
8.156
+0.139† [+0.000,+0.278]
+0.094 [−0.044,+0.233]
Gemini 2.0F
7.578
7.300
7.300
7.167
+0.206† [+0.039,+0.372]
+0.206† [+0.044,+0.372]
Llama 3.1 8B
7.656
7.344
7.378
7.322
+0.150 [−0.045,+0.344]
+0.183 [−0.017,+0.378]
Cross-model
7.758
7.525
7.467
7.447
+0.185† [+0.093,+0.275]
+0.126† [+0.037,+0.218]
Table 4: Study 2 – Critical Contrast: UNAM (Mexico, high prestige) vs. Framingham State (USA, low prestige).
Model
UNAM
Framingham St.
UNAM−FSU
Interpretation
Claude Haiku 4.5
7.233
7.011
+0.222
Prestige wins
GPT-4o-mini
8.222
8.178
+0.044
≈ Equal
Gemini 2.0 Flash
7.300
7.300
±0.000
≈ Equal
Llama 3.1 8B
7.344
7.378
−0.033
≈ Equal
Cross-model
7.525
7.467
+0.058
[−0.072,+0.186] – trend, n.s.
Table 5: Study 2 – Prestige vs. Country Effect by Domain with 95% Bootstrap CIs. † = CI entirely positive. ∗∗ = normatively unjustified (no legitimate weight under anti-discrimination principles).
Domain
Prestige effect (95% CI)
Country effect (95% CI)
Dominant
Note
Scholarship
+0.160[−0.021,+0.340]
+0.090[−0.090,+0.271]
—
Partially justified
Hiring
+0.160†[+0.035,+0.292]
+0.187†[+0.062,+0.312]
Country
Partially justified
Credit
+0.278†[+0.076,+0.479]
+0.264†[+0.062,+0.465]
Prestige
Not justified ∗∗
Health
+0.125[−0.076,+0.326]
+0.097[−0.104,+0.299]
—
Partially justified
Public Policy
+0.201†[+0.028,+0.375]
−0.007[−0.181,+0.167]
Prestige
Not justified ∗∗
Table 6: Study 3 – 2×2 Journal × Institution Prestige Cell Means. All four models, 5 domains, 3 names, 6 reps per cell. † = 95% bootstrap CI entirely positive.
Nature (hi-journal)
NCML (lo-journal)
MIT (hi-inst)
UGye (lo-inst)
MIT (hi-inst)
UGye (lo-inst)
Mean score
7.971
7.822
6.225
5.694
Journal effect (institution row)
ΔMIT=+1.746†[+1.564,+1.925]
ΔUGye=+2.128†[+1.950,+2.297]
Cross-cell main effects:
Journal (Nature−NCML)
+1.937† [+1.811,+2.062]
Institution (MIT−UGye)
+0.341† [+0.184,+0.504]
Ratio journal/institution
5.7×
Interaction (rescue effect)
−0.382 (Nature rescues UGye more than MIT)
Table 7: Study 3 – Per-Model Journal and Institution Prestige Effects with 95% CIs. † = CI entirely positive.
Model
Journal effect (95% CI)
Institution effect (95% CI)
Claude Haiku 4.5
+2.937†[+2.733,+3.139]
+0.232[−0.128,+0.594]
GPT-4o-mini
+1.322†[+1.139,+1.506]
+0.099[−0.133,+0.328]
Gemini 2.0 Flash
+2.611†[+2.372,+2.844]
+0.598†[+0.244,+0.956]
Llama 3.1 8B
+0.875†[+0.694,+1.053]
+0.430†[+0.231,+0.631]
Cross-model
+1.937†[+1.811,+2.062]
+0.341†[+0.184,+0.504]
Table 8: NBI ⟨T,I,F⟩ for Reference (Anglo-MIT) and T5 Profiles.
We investigate whether large language models (LLMs) systematically discriminate in candidate evaluations based on applicant name ethnicity and/or institutional prestige and geographic location. Three factorial experiments are reported (4,320 API calls, four LLMs, five professional domains). Study 1 (3x4 design) finds a statistically robust institution-tier gradient of +0.297 points on a 10-point scale (95% bootstrap CI: +0.175 to +0.422), while name-origin effects are negligible and non-significant (95% CI crosses zero). Study 2 (2x2 Prestige x Country design) breaks the prestige-geography confound: the prestige effect (+0.185; 95% CI: +0.093 to +0.275) exceeds the country-of-origin effect (+0.126; 95% CI: +0.037 to +0.218) by 1.5x. Study 3 (2x2 Journal x Institution design) reveals that journal prestige (Nature vs. a peripheral open-access journal) dominates institutional prestige by 5.7x: journal effect +1.937 (95% CI: +1.811 to +2.062) vs. institution effect +0.341 (95% CI: +0.184 to +0.504). A "rescue effect" is confirmed: publishing in Nature compensates for low institutional prestige more strongly for candidates from the University of Guayaquil (+2.127) than from MIT (+1.745). Results are quantified using the Neutrosophic Bias Index NBI ; the I component reveals elevated evaluation inconsistency for low-prestige profiles, an epistemic disadvantage not captured by mean-only metrics. Code and data: https://github.com/mleyvaz/geo-bias-llm