Position: AI Leaderboards Are Underserving the Global South: A Case Study from India
印度等全球南方地区其实已经有优质AI基准测试,缺的是能公正排名的独立裁判机构
这篇立场论文认为,AI排行榜之所以未能很好服务全球南方地区,根源不在于缺少数据,而在于治理机制的缺失。印度的IndicSUPERB、MILU、LAHAJA,非洲的IrokoBench,阿拉伯语的AlGhafa等高质量区域基准测试早已存在,但全球排行榜并未采用它们,也没有任何机制强制其采用。作者对82名印度AI从业者的调研显示,所有受访者都支持建立正式治理机制,其中64%更倾向于由非政府机构主导。
METAL MEDIA 解读图
印度等全球南方地区其实已经有优质AI基准测试,缺的是能公正排名的独立裁判机构
- 01MMLU、HuggingFace Open ASR排行榜等全球基准测试严重偏向欧洲语言,几乎忽略了印地语、阿拉伯语、斯瓦希里语等使用人数达数亿的语言
- 02印度的IndicSUPERB、MILU、LAHAJA,非洲的IrokoBench,阿拉伯语的AlGhafa等优质区域基准测试已经存在,但全球排行榜并不采用
- 03不少排行榜存在结构性利益冲突:同一批人既开发基准测试、又提交自家模型接受评测,却没有公开的利益冲突政策或申诉机制
- 042025年12月至2026年3月对82名印度AI从业者的调研显示,所有受访者都支持建立正式治理体系,64%倾向由非政府独立机构运营,68%倾向以信息披露而非事先排除的方式管理利益冲突
- 05作者指出,市场压力能推动修复英语相关的问题(客户投诉、合同压力),但全球南方地区缺乏这种筹码,因此真正的解法不是技术层面,而是建立治理机构、利益冲突披露制度和申诉流程
他们做了什么
- MMLU、HuggingFace Open ASR排行榜等全球基准测试严重偏向欧洲语言,几乎忽略了印地语、阿拉伯语、斯瓦希里语等使用人数达数亿的语言
- 印度的IndicSUPERB、MILU、LAHAJA,非洲的IrokoBench,阿拉伯语的AlGhafa等优质区域基准测试已经存在,但全球排行榜并不采用
- 不少排行榜存在结构性利益冲突:同一批人既开发基准测试、又提交自家模型接受评测,却没有公开的利益冲突政策或申诉机制
- 2025年12月至2026年3月对82名印度AI从业者的调研显示,所有受访者都支持建立正式治理体系,64%倾向由非政府独立机构运营,68%倾向以信息披露而非事先排除的方式管理利益冲突
- 作者指出,市场压力能推动修复英语相关的问题(客户投诉、合同压力),但全球南方地区缺乏这种筹码,因此真正的解法不是技术层面,而是建立治理机构、利益冲突披露制度和申诉流程
| Stakeholder | Impact in Multilingual AI Context |
|---|---|
| Governments | Select Automatic Speech Recognition (ASR) for citizen services using English-optimized metrics; deployed systems fail on regional dialects and code-switching111Alternating between two or more languages within a single utterance or conversation. |
| Enterprises | Choose customer service AI based on WER rankings that do not reflect local code-switching patterns (Hindi-English, Arabic-French, Spanglish) |
| Startups | Cannot benchmark regional language specialization against global models; investor due diligence relies on irrelevant metrics |
| Investors | May reference MMLU scores in due diligence; these do not predict regional language capability |
| Researchers | Publish on benchmarks that reward English optimization; local language work appears “lower performing” |
| End Users | Receive AI assistants that fail on agglutinative morphology, accents, and cultural context |
| Benchmark | Type | Coverage |
|---|---|---|
| IndicSUPERB (Javed et al., 2023a) | ASR | Multi-language evaluation across clean, noisy, telephonic conditions |
| MILU (Verma et al., 2025) | LLM | Understanding across 8 domains with Indian context |
| IndicGenBench (Singh et al., 2024) | LLM | Generation evaluation across multiple Indian languages |
| LAHAJA (Javed et al., 2024) | Dialect | Hindi regional accent evaluation |
| Vistaar (Bhogale et al., 2023) | Domain | News, education, literary, conversational contexts |
| Svarah (Javed et al., 2023b) | Accent | Indian English L2 patterns and code-mixing |
| IISc-MILE (IISc Machine Intelligence and Language Engineering Lab(2024), MILE) | Morphology | Agglutinative evaluation for Tamil and Kannada |
| Finding | Score | What it tells us |
|---|---|---|
| Overall viability | 4.06/5 | Community judges the proposal feasible |
| Endorse formal governance | 100% | No respondent chose “no governance” |
| Non-government stewardship | 64% | Nasscom + Independent NP, vs 21% govt., 9.8% academia |
| Conflict management | 68% | Disclosure/recusal over exclusion |
| Anti-gaming strategy | 56% | Dynamic evaluation, even at cost of reproducibility (5 abstentions; 60% of those who answered) |
| Evaluation method | 76% | Hybrid (LLM + human) |
| Submission cadence | 74% | Quarterly windows |
| Stakeholder | Share | Constraints |
|---|---|---|
| Academia | 30% | Max 3 seats |
| Industry | 30% | No company >1 seat |
| Government | 20% | Observer or voting |
| Civil Society | 10% | 1 seat |
| International | 10% | External perspective |
| Terms: 3 years (max 2 consecutive); Chair rotates; 2-year cooling-off period. |
| Category | Requirements |
|---|---|
| Disclosure | Annual financial interests; submitted models; funding relationships; public register |
| Recusal | Decisions affecting own interests; methodology changes; affiliated disputes |
| Prohibited | Board cannot submit models; staff cannot hold equity; no consulting with evaluated entities |
为什么重要
政府采购、企业选型、投资决策都会参考排行榜排名,如果这些排名系统性地忽略某些语言和地区,实际部署中的失败就会长期得不到纠正。这篇论文强调,当下缺的不是更多基准数据,而是一个独立、可问责的机构来汇总和评判已有的数据。
本文术语
- 排行榜(leaderboard) · 汇总AI模型基准测试分数并进行排名的平台
- 利益冲突(conflict of interest, COI) · 从高排名中获益的一方同时掌控评测过程的结构性问题
- 全球南方(Global South) · 泛指亚洲、非洲、拉丁美洲等发展中经济体较多的地区
- 词错误率(Word Error Rate, WER) · 衡量语音识别错误的指标,通过对比识别文本与参考文本计算,最初是针对英语设计的
- 古德哈特定律(Goodhart's Law) · 指一旦某个指标变成追求的目标,它就不再是一个可靠的衡量标准
论文原文摘要(英文)
This position paper argues that AI leaderboards are structurally ill-suited to serving the Global South because they lack independent governance, conflict-of-interest policies, and mechanisms for metric evolution. The barrier is not missing data; high-quality regional benchmarks already exist: IndicSUPERB, MILU, and LAHAJA for India; IrokoBench for Africa; AlGhafa for Arabic. The barrier is institutional design. Global leaderboards do not include these benchmarks, and no governance mechanism compels them to do so. Commercial pressure corrects leaderboard failures when paying customers in the Global North are affected. The Global South lacks equivalent leverage. Without governance, failures affecting Hindi, Swahili, or Arabic speakers persist indefinitely as documented but unaddressed gaps. Using India as a case study (1.4 billion people, 22 scheduled languages, high-quality benchmarks, but no trusted aggregation), we report findings from a consultation with 58 AI practitioners showing consistent preference for formal governance and disclosure-based conflict management. The solution is not more data but better institutions: regional leaderboards with independent governance from the start.
在 arXiv 阅读最新论文
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?让AI编程助手去修复真实科学软件,连最强的那个也有一半以上任务没做对
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving把稀疏注意力从论文原型变成能真正上线服务的加速方案
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents让客服AI坐席不只是拦住一个危险动作,而是把整个流程走对
- EXIMO: VLM Guided Exploration of VLA Policies不用人工遥控演示,让会说话的AI来教机械臂做新家务
- EnvHarness: Awakening Static Worlds for Agent Learning不重新搭建训练环境,而是给现有环境套一层可插拔组件,针对每个智能体的具体弱点重新塑形
- Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the Model租用AI而非拥有AI的机构,安全监管能力只剩一半
- Beyond Imitation: Filtering On-Policy Distillation by Reasoning ProgressAI模仿老师模型学习时,会误伤本来推理正确的步骤,新方法专门过滤掉这种误伤
- PersonalBench: Measuring the Authorship Gap in LLM Personalization让AI模仿某人的文风,结果发现它始终摆脱不了自己的腔调