Everything K-culture — comebacks to K-beauty, straight to your inboxGet it in your inbox

METAL MEDIA

Position: AI Leaderboards Are Underserving the Global South: A Case Study from India

arXiv:2608.181172026-08-20

India and other Global South regions already have solid AI benchmarks, but no trusted referee to rank results fairly

This position paper argues that AI leaderboards fail the Global South not because of missing data but because of weak institutional governance. High-quality regional benchmarks like IndicSUPERB, MILU, and LAHAJA for India, IrokoBench for Africa, and AlGhafa for Arabic already exist, yet global leaderboards simply don't use them, and no oversight body compels them to. A consultation with 82 Indian AI practitioners found universal support for formal governance, with 64% preferring a non-government body to run it.

METAL MEDIA explanatory visual

India and other Global South regions already have solid AI benchmarks, but no trusted referee to rank results fairly

  1. 01Global benchmarks such as MMLU and the HuggingFace Open ASR Leaderboard skew heavily toward European languages, leaving out Hindi, Arabic, Swahili, and other languages spoken by hundreds of millions of people
  2. 02Good regional benchmarks already exist (IndicSUPERB, MILU, LAHAJA for India; IrokoBench for Africa; AlGhafa for Arabic), but global leaderboards don't adopt them
  3. 03Many leaderboards let the same organizations build benchmarks, submit their own models, and judge the results, with no published conflict-of-interest policy or dispute process
  4. 04A survey of 82 Indian AI practitioners (Dec 2025-Mar 2026) found every respondent wanted formal governance; 64% preferred non-government stewardship and 68% preferred disclosure-based conflict management over outright exclusion
  5. 05The authors argue that while market pressure forces fixes for English-language failures, the Global South lacks that leverage, so the real fix has to be institutional -- governance boards, conflict-of-interest disclosure, and appeals processes -- not just more technical work
An explanatory diagram made by METAL MEDIA, not a figure supplied by the paper's authors.

What they did

  1. Global benchmarks such as MMLU and the HuggingFace Open ASR Leaderboard skew heavily toward European languages, leaving out Hindi, Arabic, Swahili, and other languages spoken by hundreds of millions of people
  2. Good regional benchmarks already exist (IndicSUPERB, MILU, LAHAJA for India; IrokoBench for Africa; AlGhafa for Arabic), but global leaderboards don't adopt them
  3. Many leaderboards let the same organizations build benchmarks, submit their own models, and judge the results, with no published conflict-of-interest policy or dispute process
  4. A survey of 82 Indian AI practitioners (Dec 2025-Mar 2026) found every respondent wanted formal governance; 64% preferred non-government stewardship and 68% preferred disclosure-based conflict management over outright exclusion
  5. The authors argue that while market pressure forces fixes for English-language failures, the Global South lacks that leverage, so the real fix has to be institutional -- governance boards, conflict-of-interest disclosure, and appeals processes -- not just more technical work
Table 1: Stakeholder Impact of Unreliable Leaderboards
StakeholderImpact in Multilingual AI Context
GovernmentsSelect Automatic Speech Recognition (ASR) for citizen services using English-optimized metrics; deployed systems fail on regional dialects and code-switching111Alternating between two or more languages within a single utterance or conversation.
EnterprisesChoose customer service AI based on WER rankings that do not reflect local code-switching patterns (Hindi-English, Arabic-French, Spanglish)
StartupsCannot benchmark regional language specialization against global models; investor due diligence relies on irrelevant metrics
InvestorsMay reference MMLU scores in due diligence; these do not predict regional language capability
ResearchersPublish on benchmarks that reward English optimization; local language work appears “lower performing”
End UsersReceive AI assistants that fail on agglutinative morphology, accents, and cultural context
Table 2: Existing Indic AI Benchmarks (non-exhaustive)
BenchmarkTypeCoverage
IndicSUPERB (Javed et al., 2023a)ASRMulti-language evaluation across clean, noisy, telephonic conditions
MILU (Verma et al., 2025)LLMUnderstanding across 8 domains with Indian context
IndicGenBench (Singh et al., 2024)LLMGeneration evaluation across multiple Indian languages
LAHAJA (Javed et al., 2024)DialectHindi regional accent evaluation
Vistaar (Bhogale et al., 2023)DomainNews, education, literary, conversational contexts
Svarah (Javed et al., 2023b)AccentIndian English L2 patterns and code-mixing
IISc-MILE (IISc Machine Intelligence and Language Engineering Lab(2024), MILE)MorphologyAgglutinative evaluation for Tamil and Kannada
Table 3: Stakeholder Consultation Key Findings (n=82, Dec 2025 – Mar 2026)
FindingScoreWhat it tells us
Overall viability4.06/5Community judges the proposal feasible
Endorse formal governance100%No respondent chose “no governance”
Non-government stewardship64%Nasscom + Independent NP, vs 21% govt., 9.8% academia
Conflict management68%Disclosure/recusal over exclusion
Anti-gaming strategy56%Dynamic evaluation, even at cost of reproducibility (5 abstentions; 60% of those who answered)
Evaluation method76%Hybrid (LLM + human)
Submission cadence74%Quarterly windows
Table 4: Board Composition and Terms
StakeholderShareConstraints
Academia30%Max 3 seats
Industry30%No company >1 seat
Government20%Observer or voting
Civil Society10%1 seat
International10%External perspective
Terms: 3 years (max 2 consecutive); Chair rotates; 2-year cooling-off period.
Table 5: Conflict of Interest Policy
CategoryRequirements
DisclosureAnnual financial interests; submitted models; funding relationships; public register
RecusalDecisions affecting own interests; methodology changes; affiliated disputes
ProhibitedBoard cannot submit models; staff cannot hold equity; no consulting with evaluated entities

Why it matters

Governments, companies, and investors all lean on leaderboard rankings for procurement and decision-making, so if those rankings systematically ignore certain languages and regions, real deployment failures go uncorrected. The paper's point is that the missing piece isn't more benchmark data, it's an independent, accountable body to aggregate and rank what already exists.

Terms in this paper

  • leaderboard · a platform that collects benchmark scores from AI models and ranks them
  • conflict of interest (COI) · a situation where the same party that benefits from high rankings also controls how evaluation is done
  • Global South · a general term for regions such as Asia, Africa, and Latin America that include many developing economies
  • Word Error Rate (WER) · a metric that measures speech recognition errors by comparing output text to a reference, originally designed around English
  • Goodhart's Law · the principle that once a measure becomes a target, it stops being a reliable measure

Original abstract (English)

This position paper argues that AI leaderboards are structurally ill-suited to serving the Global South because they lack independent governance, conflict-of-interest policies, and mechanisms for metric evolution. The barrier is not missing data; high-quality regional benchmarks already exist: IndicSUPERB, MILU, and LAHAJA for India; IrokoBench for Africa; AlGhafa for Arabic. The barrier is institutional design. Global leaderboards do not include these benchmarks, and no governance mechanism compels them to do so. Commercial pressure corrects leaderboard failures when paying customers in the Global North are affected. The Global South lacks equivalent leverage. Without governance, failures affecting Hindi, Swahili, or Arabic speakers persist indefinitely as documented but unaddressed gaps. Using India as a case study (1.4 billion people, 22 scheduled languages, high-quality benchmarks, but no trusted aggregation), we report findings from a consultation with 58 AI practitioners showing consistent preference for formal governance and disclosure-based conflict management. The solution is not more data but better institutions: regional leaderboards with independent governance from the start.

Authors · Sourav Banerjee, Saikat Saha

Read on arXiv

Latest papers

All papers →

Latest from METAL MEDIA