Everything K-culture — comebacks to K-beauty, straight to your inboxGet it in your inbox

METAL MEDIA

Institutional Prestige as Geographic Bias in Large Language Models: Evidence from Three Factorial Experiments with Bootstrap Confidence Intervals

arXiv:2608.181072026-08-20

In AI candidate scoring, school prestige and where you publish matter far more than your name

Testing four large language models as automated evaluators across scholarship, hiring, credit, health and policy scenarios, researchers found that the ethnic signal carried by a candidate's name barely moved scores, but institutional prestige and journal prestige did, strongly. A single publication in Nature swung scores 5.7 times more than attending a top-tier university did. The team introduced a new Neutrosophic Bias Index (NBI) showing that low-prestige candidates are hurt twice over: lower average scores and more inconsistent evaluations.

METAL MEDIA explanatory visual

In AI candidate scoring, school prestige and where you publish matter far more than your name

  1. 01Across 4,320 API calls to Claude Haiku 4.5, GPT-4o-mini, Gemini 2.0 Flash, and Llama 3.1 8B, the study ran three factorial experiments varying applicant name origin (Anglo/Latino/Arabic), institution prestige tier, country, and journal prestige across five professional domains
  2. 02Name-origin effects on 10-point scores were statistically non-significant (±0.094, confidence interval crosses zero), but institution tier produced a robust gradient of +0.297 points from top-ranked (MIT-level) to unranked institutions
  3. 03Separating institutional prestige (+0.185) from country of origin (developed vs developing, +0.126) showed prestige mattered 1.5 times more than country, and a prestigious Latin American university (UNAM) was not rated below an unranked US university (Framingham State), ruling out pure country bias
  4. 04Crossing journal prestige (Nature vs. a peripheral open-access journal) with institutional prestige revealed the journal effect (+1.937) was 5.7 times larger than the institution effect (+0.341); publishing in Nature boosted scores for candidates from a low-prestige university (Universidad de Guayaquil) even more than for MIT candidates, a 'rescue effect'
  5. 05The newly introduced Neutrosophic Bias Index (NBI) showed low-prestige candidates face a compound disadvantage: not just lower average scores but also higher evaluation inconsistency, a form of epistemic disadvantage that mean-score-only audits miss
An explanatory diagram made by METAL MEDIA, not a figure supplied by the paper's authors.

What they did

  1. Across 4,320 API calls to Claude Haiku 4.5, GPT-4o-mini, Gemini 2.0 Flash, and Llama 3.1 8B, the study ran three factorial experiments varying applicant name origin (Anglo/Latino/Arabic), institution prestige tier, country, and journal prestige across five professional domains
  2. Name-origin effects on 10-point scores were statistically non-significant (±0.094, confidence interval crosses zero), but institution tier produced a robust gradient of +0.297 points from top-ranked (MIT-level) to unranked institutions
  3. Separating institutional prestige (+0.185) from country of origin (developed vs developing, +0.126) showed prestige mattered 1.5 times more than country, and a prestigious Latin American university (UNAM) was not rated below an unranked US university (Framingham State), ruling out pure country bias
  4. Crossing journal prestige (Nature vs. a peripheral open-access journal) with institutional prestige revealed the journal effect (+1.937) was 5.7 times larger than the institution effect (+0.341); publishing in Nature boosted scores for candidates from a low-prestige university (Universidad de Guayaquil) even more than for MIT candidates, a 'rescue effect'
  5. The newly introduced Neutrosophic Bias Index (NBI) showed low-prestige candidates face a compound disadvantage: not just lower average scores but also higher evaluation inconsistency, a form of epistemic disadvantage that mean-score-only audits miss
Figure 1: Study 1 – Institution-Tier Gradient by Model with 95% Bootstrap CI (cross-model).
Figure 1: Study 1 – Institution-Tier Gradient by Model with 95% Bootstrap CI (cross-model).
Table 1: Study 1 – Mean Score by Institution Tier. Cross-model CI from 10,000 bootstrap iterations. ✓ = 95% CI entirely positive.
ModelT1 MITT2 UChileT3 UNALT5 UGyeGradient T1−T595% CI
Claude Haiku 4.57.4337.2337.2227.133+0.300[+0.111,+0.478]
GPT-4o-mini8.3788.2338.2118.156+0.222[+0.022,+0.422]
Gemini 2.0 Flash7.5567.3787.3117.189+0.367[+0.133,+0.611]
Llama 3.1 8B7.6787.4447.589∗7.378+0.300[+0.044,+0.556]
Cross-model7.7617.5727.5837.464+0.297[+0.175,+0.422] ✓
Figure 2: Study 2 – 2×2 Prestige × Country Cell Means by Model.
Figure 2: Study 2 – 2×2 Prestige × Country Cell Means by Model.
Table 2: Study 1 – Mean Score by Name Origin. Bootstrap CIs cross zero for all contrasts (⇒ non-significant).
ModelAngloLatinoArabicMax gapSignificant?
Claude Haiku 4.57.2087.2337.3250.117No
GPT-4o-mini8.2178.2508.2670.050No
Gemini 2.0 Flash7.2677.4177.3920.150No
Llama 3.1 8B7.4677.5507.5500.083No
Cross-model7.5407.6127.6330.094No (CI ∋ 0)
Figure 3: Study 3 – 2×2 Journal × Institution Prestige. The “rescue effect” cell (UGye + Nature) shows the largest journal premium (Δ=+2.13), indicating that Nature publication compensates for low institutional prestige more than for high institutional prestige.
Figure 3: Study 3 – 2×2 Journal × Institution Prestige. The “rescue effect” cell (UGye + Nature) shows the largest journal premium (Δ=+2.13), indicating that Nature publication compensates for low institutional prestige more than for high institutional prestige.
Table 3: Study 2 – 2×2 Cell Means and Factorial Effects with 95% Bootstrap CIs. † = CI entirely positive (significant). Dev = Developed; Dvlp = Developing.
ModelMITUNAMFSUUGyePrestige (95% CI)Country (95% CI)
(Hi,Dev)(Hi,Dvlp)(Lo,Dev)(Lo,Dvlp)
Haiku 4.57.4117.2337.0117.144+0.244† [+0.106,+0.383]+0.022 [−0.117,+0.167]
GPT-4o-mini8.3898.2228.1788.156+0.139† [+0.000,+0.278]+0.094 [−0.044,+0.233]
Gemini 2.0F7.5787.3007.3007.167+0.206† [+0.039,+0.372]+0.206† [+0.044,+0.372]
Llama 3.1 8B7.6567.3447.3787.322+0.150 [−0.045,+0.344]+0.183 [−0.017,+0.378]
Cross-model7.7587.5257.4677.447+0.185† [+0.093,+0.275]+0.126† [+0.037,+0.218]
Table 4: Study 2 – Critical Contrast: UNAM (Mexico, high prestige) vs. Framingham State (USA, low prestige).
ModelUNAMFramingham St.UNAM−FSUInterpretation
Claude Haiku 4.57.2337.011+0.222Prestige wins
GPT-4o-mini8.2228.178+0.044≈ Equal
Gemini 2.0 Flash7.3007.300±0.000≈ Equal
Llama 3.1 8B7.3447.378−0.033≈ Equal
Cross-model7.5257.467+0.058[−0.072,+0.186] – trend, n.s.
Table 5: Study 2 – Prestige vs. Country Effect by Domain with 95% Bootstrap CIs. † = CI entirely positive. ∗⁣∗ = normatively unjustified (no legitimate weight under anti-discrimination principles).
DomainPrestige effect (95% CI)Country effect (95% CI)DominantNote
Scholarship+0.160​[−0.021,+0.340]+0.090​[−0.090,+0.271]Partially justified
Hiring+0.160†​[+0.035,+0.292]+0.187†​[+0.062,+0.312]CountryPartially justified
Credit+0.278†​[+0.076,+0.479]+0.264†​[+0.062,+0.465]PrestigeNot justified ∗⁣∗
Health+0.125​[−0.076,+0.326]+0.097​[−0.104,+0.299]Partially justified
Public Policy+0.201†​[+0.028,+0.375]−0.007​[−0.181,+0.167]PrestigeNot justified ∗⁣∗
Table 6: Study 3 – 2×2 Journal × Institution Prestige Cell Means. All four models, 5 domains, 3 names, 6 reps per cell. † = 95% bootstrap CI entirely positive.
Nature (hi-journal)NCML (lo-journal)
MIT (hi-inst)UGye (lo-inst)MIT (hi-inst)UGye (lo-inst)
Mean score7.9717.8226.2255.694
Journal effect (institution row)ΔMIT=+1.746†​[+1.564,+1.925]ΔUGye=+2.128†​[+1.950,+2.297]
Cross-cell main effects:
Journal (Nature−NCML)+1.937† [+1.811,+2.062]
Institution (MIT−UGye)+0.341† [+0.184,+0.504]
Ratio journal/institution5.7×
Interaction (rescue effect)−0.382 (Nature rescues UGye more than MIT)
Table 7: Study 3 – Per-Model Journal and Institution Prestige Effects with 95% CIs. † = CI entirely positive.
ModelJournal effect (95% CI)Institution effect (95% CI)
Claude Haiku 4.5+2.937†​[+2.733,+3.139]+0.232​[−0.128,+0.594]
GPT-4o-mini+1.322†​[+1.139,+1.506]+0.099​[−0.133,+0.328]
Gemini 2.0 Flash+2.611†​[+2.372,+2.844]+0.598†​[+0.244,+0.956]
Llama 3.1 8B+0.875†​[+0.694,+1.053]+0.430†​[+0.231,+0.631]
Cross-model+1.937†​[+1.811,+2.062]+0.341†​[+0.184,+0.504]
Table 8: NBI ⟨T,I,F⟩ for Reference (Anglo-MIT) and T5 Profiles.
ModelProfileTIFInterpretation
Haiku 4.5Anglo T1 (ref)0.7400.0790.000Reference
Anglo T50.7100.1180.030Prestige penalty
Latino T50.7070.1090.033Prestige penalty
Arabic T50.7230.1090.017Prestige penalty
GPT-4o-miniAnglo T1 (ref)0.8280.1260.000Reference
Anglo T50.8110.1260.017Prestige penalty
Latino T50.8090.1250.020Prestige penalty
Arabic T50.8110.1260.017Prestige penalty
Gemini 2.0FAnglo T1 (ref)0.7530.1390.000Reference
Anglo T50.7050.1260.046Prestige penalty
Latino T50.7230.1080.033Prestige penalty
Arabic T50.7270.1040.027Prestige penalty
Llama 3.1 8BAnglo T1 (ref)0.7800.1150.000Reference
Anglo T50.7220.1870.057Strongest bias
Latino T50.7470.1530.033Prestige penalty
Arabic T50.7430.1340.037Prestige penalty

Why it matters

As LLMs get deployed for scholarship, hiring, credit and grant decisions, checking only for name-based bias can give a false sense of fairness while institutional and publication-venue bias go undetected. This work empirically shows how researchers and applicants from lower-prestige institutions or the Global South could be systematically disadvantaged by AI evaluators.

Terms in this paper

  • bootstrap confidence interval · a statistical range estimated by repeatedly resampling the data to gauge how reliable a measured effect is
  • Neutrosophic Bias Index (NBI) · a new metric proposed in this paper expressing bias as three values, truth, indeterminacy, and falsity, capturing both average score penalty and evaluation inconsistency
  • factorial design · an experimental setup that varies multiple factors (like name, institution tier, country, journal) at once to isolate each factor's independent effect
  • rescue effect · the finding that publishing in a prestigious journal compensates for low institutional prestige, more strongly than for candidates from high-prestige institutions

Original abstract (English)

We investigate whether large language models (LLMs) systematically discriminate in candidate evaluations based on applicant name ethnicity and/or institutional prestige and geographic location. Three factorial experiments are reported (4,320 API calls, four LLMs, five professional domains). Study 1 (3x4 design) finds a statistically robust institution-tier gradient of +0.297 points on a 10-point scale (95% bootstrap CI: +0.175 to +0.422), while name-origin effects are negligible and non-significant (95% CI crosses zero). Study 2 (2x2 Prestige x Country design) breaks the prestige-geography confound: the prestige effect (+0.185; 95% CI: +0.093 to +0.275) exceeds the country-of-origin effect (+0.126; 95% CI: +0.037 to +0.218) by 1.5x. Study 3 (2x2 Journal x Institution design) reveals that journal prestige (Nature vs. a peripheral open-access journal) dominates institutional prestige by 5.7x: journal effect +1.937 (95% CI: +1.811 to +2.062) vs. institution effect +0.341 (95% CI: +0.184 to +0.504). A "rescue effect" is confirmed: publishing in Nature compensates for low institutional prestige more strongly for candidates from the University of Guayaquil (+2.127) than from MIT (+1.745). Results are quantified using the Neutrosophic Bias Index NBI ; the I component reveals elevated evaluation inconsistency for low-prestige profiles, an epistemic disadvantage not captured by mean-only metrics. Code and data: https://github.com/mleyvaz/geo-bias-llm

Authors · Maikel Leyva-Vazquez, Florentin Smarandache

Read on arXiv

Latest papers

All papers →

Latest from METAL MEDIA

Figures: Maikel Leyva-Vazquez et al., arXiv:2608.18107, CC BY 4.0