Position: AI Leaderboards Are Underserving the Global South: A Case Study from India
India and other Global South regions already have solid AI benchmarks, but no trusted referee to rank results fairly
This position paper argues that AI leaderboards fail the Global South not because of missing data but because of weak institutional governance. High-quality regional benchmarks like IndicSUPERB, MILU, and LAHAJA for India, IrokoBench for Africa, and AlGhafa for Arabic already exist, yet global leaderboards simply don't use them, and no oversight body compels them to. A consultation with 82 Indian AI practitioners found universal support for formal governance, with 64% preferring a non-government body to run it.
METAL MEDIA explanatory visual
India and other Global South regions already have solid AI benchmarks, but no trusted referee to rank results fairly
- 01Global benchmarks such as MMLU and the HuggingFace Open ASR Leaderboard skew heavily toward European languages, leaving out Hindi, Arabic, Swahili, and other languages spoken by hundreds of millions of people
- 02Good regional benchmarks already exist (IndicSUPERB, MILU, LAHAJA for India; IrokoBench for Africa; AlGhafa for Arabic), but global leaderboards don't adopt them
- 03Many leaderboards let the same organizations build benchmarks, submit their own models, and judge the results, with no published conflict-of-interest policy or dispute process
- 04A survey of 82 Indian AI practitioners (Dec 2025-Mar 2026) found every respondent wanted formal governance; 64% preferred non-government stewardship and 68% preferred disclosure-based conflict management over outright exclusion
- 05The authors argue that while market pressure forces fixes for English-language failures, the Global South lacks that leverage, so the real fix has to be institutional -- governance boards, conflict-of-interest disclosure, and appeals processes -- not just more technical work
What they did
- Global benchmarks such as MMLU and the HuggingFace Open ASR Leaderboard skew heavily toward European languages, leaving out Hindi, Arabic, Swahili, and other languages spoken by hundreds of millions of people
- Good regional benchmarks already exist (IndicSUPERB, MILU, LAHAJA for India; IrokoBench for Africa; AlGhafa for Arabic), but global leaderboards don't adopt them
- Many leaderboards let the same organizations build benchmarks, submit their own models, and judge the results, with no published conflict-of-interest policy or dispute process
- A survey of 82 Indian AI practitioners (Dec 2025-Mar 2026) found every respondent wanted formal governance; 64% preferred non-government stewardship and 68% preferred disclosure-based conflict management over outright exclusion
- The authors argue that while market pressure forces fixes for English-language failures, the Global South lacks that leverage, so the real fix has to be institutional -- governance boards, conflict-of-interest disclosure, and appeals processes -- not just more technical work
| Stakeholder | Impact in Multilingual AI Context |
|---|---|
| Governments | Select Automatic Speech Recognition (ASR) for citizen services using English-optimized metrics; deployed systems fail on regional dialects and code-switching111Alternating between two or more languages within a single utterance or conversation. |
| Enterprises | Choose customer service AI based on WER rankings that do not reflect local code-switching patterns (Hindi-English, Arabic-French, Spanglish) |
| Startups | Cannot benchmark regional language specialization against global models; investor due diligence relies on irrelevant metrics |
| Investors | May reference MMLU scores in due diligence; these do not predict regional language capability |
| Researchers | Publish on benchmarks that reward English optimization; local language work appears “lower performing” |
| End Users | Receive AI assistants that fail on agglutinative morphology, accents, and cultural context |
| Benchmark | Type | Coverage |
|---|---|---|
| IndicSUPERB (Javed et al., 2023a) | ASR | Multi-language evaluation across clean, noisy, telephonic conditions |
| MILU (Verma et al., 2025) | LLM | Understanding across 8 domains with Indian context |
| IndicGenBench (Singh et al., 2024) | LLM | Generation evaluation across multiple Indian languages |
| LAHAJA (Javed et al., 2024) | Dialect | Hindi regional accent evaluation |
| Vistaar (Bhogale et al., 2023) | Domain | News, education, literary, conversational contexts |
| Svarah (Javed et al., 2023b) | Accent | Indian English L2 patterns and code-mixing |
| IISc-MILE (IISc Machine Intelligence and Language Engineering Lab(2024), MILE) | Morphology | Agglutinative evaluation for Tamil and Kannada |
| Finding | Score | What it tells us |
|---|---|---|
| Overall viability | 4.06/5 | Community judges the proposal feasible |
| Endorse formal governance | 100% | No respondent chose “no governance” |
| Non-government stewardship | 64% | Nasscom + Independent NP, vs 21% govt., 9.8% academia |
| Conflict management | 68% | Disclosure/recusal over exclusion |
| Anti-gaming strategy | 56% | Dynamic evaluation, even at cost of reproducibility (5 abstentions; 60% of those who answered) |
| Evaluation method | 76% | Hybrid (LLM + human) |
| Submission cadence | 74% | Quarterly windows |
| Stakeholder | Share | Constraints |
|---|---|---|
| Academia | 30% | Max 3 seats |
| Industry | 30% | No company >1 seat |
| Government | 20% | Observer or voting |
| Civil Society | 10% | 1 seat |
| International | 10% | External perspective |
| Terms: 3 years (max 2 consecutive); Chair rotates; 2-year cooling-off period. |
| Category | Requirements |
|---|---|
| Disclosure | Annual financial interests; submitted models; funding relationships; public register |
| Recusal | Decisions affecting own interests; methodology changes; affiliated disputes |
| Prohibited | Board cannot submit models; staff cannot hold equity; no consulting with evaluated entities |
Why it matters
Governments, companies, and investors all lean on leaderboard rankings for procurement and decision-making, so if those rankings systematically ignore certain languages and regions, real deployment failures go uncorrected. The paper's point is that the missing piece isn't more benchmark data, it's an independent, accountable body to aggregate and rank what already exists.
Terms in this paper
- leaderboard · a platform that collects benchmark scores from AI models and ranks them
- conflict of interest (COI) · a situation where the same party that benefits from high rankings also controls how evaluation is done
- Global South · a general term for regions such as Asia, Africa, and Latin America that include many developing economies
- Word Error Rate (WER) · a metric that measures speech recognition errors by comparing output text to a reference, originally designed around English
- Goodhart's Law · the principle that once a measure becomes a target, it stops being a reliable measure
Original abstract (English)
This position paper argues that AI leaderboards are structurally ill-suited to serving the Global South because they lack independent governance, conflict-of-interest policies, and mechanisms for metric evolution. The barrier is not missing data; high-quality regional benchmarks already exist: IndicSUPERB, MILU, and LAHAJA for India; IrokoBench for Africa; AlGhafa for Arabic. The barrier is institutional design. Global leaderboards do not include these benchmarks, and no governance mechanism compels them to do so. Commercial pressure corrects leaderboard failures when paying customers in the Global North are affected. The Global South lacks equivalent leverage. Without governance, failures affecting Hindi, Swahili, or Arabic speakers persist indefinitely as documented but unaddressed gaps. Using India as a case study (1.4 billion people, 22 scheduled languages, high-quality benchmarks, but no trusted aggregation), we report findings from a consultation with 58 AI practitioners showing consistent preference for formal governance and disclosure-based conflict management. The solution is not more data but better institutions: regional leaderboards with independent governance from the start.
Read on arXivLatest papers
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?AI coding agents were tested on fixing real scientific software, and even the best one failed more than half the time
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM ServingMaking sparse attention fast enough and accurate enough for real LLM serving, not just papers
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM AgentsMaking customer-service AI agents follow the whole procedure, not just avoid one bad action
- EXIMO: VLM Guided Exploration of VLA PoliciesTeaching a robot new chores without human teleoperation, by letting a chatty AI supervise it
- EnvHarness: Awakening Static Worlds for Agent LearningInstead of building new training worlds from scratch, this work adds a plug-in layer that reshapes existing ones around each agent's actual weaknesses
- Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the ModelCompanies that rent AI instead of owning it can only do half of AI safety oversight
- Beyond Imitation: Filtering On-Policy Distillation by Reasoning ProgressA fix for AI models that get penalized by their teacher even when they're reasoning correctly
- PersonalBench: Measuring the Authorship Gap in LLM PersonalizationAI can be prompted to write 'like someone,' but its own voice never fully disappears