NE-BERT: A Multilingual Language Model for Nine Northeast Indian Languages
为印度东北部9种少数民族语言从零训练的语言模型,理解能力反而超过大型多语言模型
隶属于MWire Labs的研究者用约830万个句子训练出名为NE-BERT的语言模型,涵盖印度东北部9种本土语言以及印地语、英语。针对句子数量极少的语言,例如只有1002句的Pnar语,团队采用了加权采样和定制分词器(把文本切成模型可处理片段的工具)来解决词汇被切碎的问题。结果显示,NE-BERT在衡量语言模型预测能力的困惑度指标上,平均分别比IndicBERT-V2和MuRIL好15.97倍和7.64倍。
METAL MEDIA 解读图
为印度东北部9种少数民族语言从零训练的语言模型,理解能力反而超过大型多语言模型
- 01印度东北部拥有超过200种语言,但现有多语言模型几乎没有覆盖这一地区
- 02团队收集了Assamese、Garo、Khasi、Meitei、Mizo、Naga、Nyishi、Pnar、Kokborok九种语言加上印地语和英语共约830万句,并在训练分词器时对Pnar(1002句)、Kokborok(2463句)这类极度稀缺语言进行100倍上采样,防止词汇被拆成无意义的字符碎片
- 03没有采用常见的字节对编码(BPE)方法,而是选用SentencePiece Unigram方式,构建了一个含5万368个词元的定制分词器,其切词效率比mBERT高1.50倍
- 04在困惑度指标(数值越低表示预测越好)上,NE-BERT在9种东北部语言上平均达到2.21,远优于IndicBERT-V2的35.29和MuRIL的16.88,在Pnar、Nyishi等极低资源语言上差距尤其明显,基线模型几乎失效
- 05在词性标注这一实际下游任务上,NE-BERT在Khasi、Mizo、Nagamese三种语言上平均准确率达82.4%,比mBERT高9.1个百分点,比IndicBERT-V2高23.2个百分点,而整个模型训练仅用单张A40 GPU耗时约17小时,成本7.31美元
他们做了什么
- 印度东北部拥有超过200种语言,但现有多语言模型几乎没有覆盖这一地区
- 团队收集了Assamese、Garo、Khasi、Meitei、Mizo、Naga、Nyishi、Pnar、Kokborok九种语言加上印地语和英语共约830万句,并在训练分词器时对Pnar(1002句)、Kokborok(2463句)这类极度稀缺语言进行100倍上采样,防止词汇被拆成无意义的字符碎片
- 没有采用常见的字节对编码(BPE)方法,而是选用SentencePiece Unigram方式,构建了一个含5万368个词元的定制分词器,其切词效率比mBERT高1.50倍
- 在困惑度指标(数值越低表示预测越好)上,NE-BERT在9种东北部语言上平均达到2.21,远优于IndicBERT-V2的35.29和MuRIL的16.88,在Pnar、Nyishi等极低资源语言上差距尤其明显,基线模型几乎失效
- 在词性标注这一实际下游任务上,NE-BERT在Khasi、Mizo、Nagamese三种语言上平均准确率达82.4%,比mBERT高9.1个百分点,比IndicBERT-V2高23.2个百分点,而整个模型训练仅用单张A40 GPU耗时约17小时,成本7.31美元

| Language | ISO | Sentences | Tokens | Virtual Count | Weight | Family | Source |
|---|---|---|---|---|---|---|---|
| Anchor Languages | |||||||
| Hindi | hin | 3,404,007 | — | 170,200 | 0.05× | Indo-Aryan | HF Datasets |
| English | eng | 500,000 | — | 100,000 | 0.2× | Germanic | HF Datasets |
| Northeast Indian Languages | |||||||
| Meitei | mni | 1,354,323 | 42,504,181 | 1,354,323 | 1.0× | Sino-Tibetan | Curated |
| Assamese | asm | 1,000,000 | 38,652,391 | 1,000,000 | 1.0× | Indo-Aryan | Curated |
| Khasi | kha | 1,000,000 | 17,472,606 | 1,000,000 | 1.0× | Austroasiatic | Curated |
| Mizo | lus | 1,000,000 | 26,774,164 | 1,000,000 | 1.0× | Sino-Tibetan | Curated |
| Nyishi | njz | 55,870 | 560,374 | 1,117,400 | 20.0× | Sino-Tibetan | WMT 2025 |
| Naga | nag | 13,918 | 508,980 | 278,360 | 20.0× | Sino-Tibetan | Curated |
| Garo | grt | 10,817 | 243,251 | 216,340 | 20.0× | Sino-Tibetan | Curated |
| Kokborok | trp | 2,463 | 89,851 | 246,300 | 100.0× | Sino-Tibetan | WMT 2025 |
| Pnar | pbv | 1,002 | 52,144 | 100,200 | 100.0× | Austroasiatic | Curated |
| NE Total | 4,438,393 | 126,857,942 | |||||
| Overall Total | 8,342,400 | — | 6,583,123 |
| Model | Params | Layers | Hidden | Heads |
|---|---|---|---|---|
| NE-BERT | 149M | 22 | 768 | 12 |
| mBERT | 110M | 12 | 768 | 12 |
| IndicBERT-V2 | 237M | 12 | 1024 | 16 |
| MuRIL | 236M | 24 | 1024 | 16 |
| Language | NE-BERT | IB-V2 | MuRIL | mBERT |
|---|---|---|---|---|
| Assamese | 1.76 | 9.01 | 5.62 | 1.65 |
| Meitei | 1.89 | 3.77 | 3.22 | 1.44 |
| Khasi | 1.27 | 2.40 | 1.93 | 1.36 |
| Mizo | 1.90 | 8.95 | 7.00 | 2.44 |
| Garo | 2.64 | 26.37 | 15.88 | 3.64 |
| Kokborok | 1.72 | 9.15 | 5.41 | 2.23 |
| Pnar | 2.92 | 66.92 | 39.65 | 5.30 |
| Naga | 1.49 | 3.83 | 3.39 | 1.81 |
| Nyishi | 4.33 | 187.20 | 75.80 | 6.05 |
| English | 1.55 | 14.81 | 5.57 | 2.64 |
| Hindi | 1.43 | 10.08 | 5.75 | 1.79 |
| NE Avg. | 2.21 | 35.29 | 16.88 | 2.77 |
| Overall | 2.17 | 31.14 | 15.38 | 2.76 |
| Language | NE-B | IB-V2 | MuRIL | mB |
|---|---|---|---|---|
| Assamese | 1.63 | 1.61 | 1.72 | 3.80 |
| Meitei | 1.52 | 2.77 | 3.01 | 4.17 |
| Khasi | 1.31 | 1.77 | 1.78 | 1.78 |
| Mizo | 1.36 | 1.73 | 1.79 | 1.78 |
| Garo | 2.37 | 2.84 | 2.80 | 2.89 |
| Kokborok | 2.07 | 2.07 | 2.22 | 2.19 |
| Pnar | 1.61 | 1.69 | 1.69 | 1.67 |
| Naga | 1.56 | 1.93 | 1.77 | 1.97 |
| Nyishi | 1.67 | 2.34 | 2.44 | 2.36 |
| NE Avg. | 1.68 | 2.08 | 2.14 | 2.51 |
| Language | NE-B | IB-V2 | MuR | mB |
|---|---|---|---|---|
| Assamese | 0.230 | 0.880 | 0.731 | 0.442 |
| Meitei | 0.220 | 0.815 | 0.778 | 0.331 |
| Khasi | 0.092 | 0.445 | 0.336 | 0.158 |
| Mizo | 0.272 | 1.163 | 1.063 | 0.486 |
| Garo | 0.503 | 1.997 | 1.663 | 0.800 |
| Kokborok | 0.277 | 1.136 | 0.926 | 0.433 |
| Pnar | 0.680 | 2.795 | 2.446 | 1.092 |
| Naga | 0.171 | 0.709 | 0.594 | 0.321 |
| Nyishi | 0.789 | 3.763 | 3.231 | 1.306 |
| English | 0.190 | 0.862 | 0.558 | 0.328 |
| Hindi | 0.192 | 0.863 | 0.659 | 0.345 |
| NE Avg. | 0.359 | 1.545 | 1.307 | 0.597 |
| Overall | 0.347 | 1.497 | 1.271 | 0.590 |
| Language | NE-BERT | mBERT | IB-V2 | MuRIL |
|---|---|---|---|---|
| Khasi | 87.7 | 82.3 | 80.1 | 61.2 |
| Mizo | 73.2 | 66.3 | 55.7 | 43.1 |
| Nagamese | 86.4 | 71.3 | 41.7 | 44.8 |
| Average | 82.4 | 73.3 | 59.2 | 49.7 |
为什么重要
这项工作说明,对于几乎没有数字化资料的语言,精心设计数据配比和分词方式比单纯扩大模型规模更关键,从而能以极低成本训练出有效的语言模型。团队以CC-BY-4.0协议开源了模型、训练代码和数据集,为资源匮乏的语言社区提供了可直接使用的起点。
本文术语
- 困惑度(Perplexity) · 衡量语言模型预测文本能力的指标,数值越低说明预测越准确
- 分词效率(fertility) · 分词器把一个单词平均拆成几个片段的数值,越低说明分词越高效
- SentencePiece Unigram · 一种基于概率的子词切分方法,相比贪婪式方法更能保留结构复杂的单词
- 加权采样 · 人为提高稀缺语言在训练数据中的比重,以缓解数据不平衡问题
- 掩码语言建模(MLM) · 一种训练方式,遮住句子的部分内容,让模型学习预测被遮住的词
论文原文摘要(英文)
Large pretrained language models have demonstrated remarkable capabilities across diverse languages, yet critically underrepresented low-resource languages remain marginalized. We present NE-BERT, a domain-specific multilingual encoder model trained on approximately 8.3 million sentences spanning 9 Northeast Indian languages and 2 anchor languages (Hindi, English), a linguistically diverse region with minimal representation in existing multilingual models. By employing weighted data sampling and a custom SentencePiece Unigram tokenizer, NE-BERT outperforms IndicBERT-V2 and MuRIL across all 9 Northeast Indian languages, achieving 15.97X and 7.64X lower average perplexity respectively, with 1.50X better tokenization fertility than mBERT. We address critical vocabulary fragmentation issues in extremely low-resource languages such as Pnar (1,002 sentences) and Kokborok (2,463 sentences) through aggressive upsampling strategies. Downstream evaluation on part-of-speech tagging validates practical utility on three Northeast Indian languages. We release NE-BERT, test sets, and training corpus under CC-BY-4.0 to support NLP research and digital inclusion for Northeast Indian communities.
在 arXiv 阅读最新论文
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?让AI编程助手去修复真实科学软件,连最强的那个也有一半以上任务没做对
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving把稀疏注意力从论文原型变成能真正上线服务的加速方案
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents让客服AI坐席不只是拦住一个危险动作,而是把整个流程走对
- EXIMO: VLM Guided Exploration of VLA Policies不用人工遥控演示,让会说话的AI来教机械臂做新家务
- EnvHarness: Awakening Static Worlds for Agent Learning不重新搭建训练环境,而是给现有环境套一层可插拔组件,针对每个智能体的具体弱点重新塑形
- Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the Model租用AI而非拥有AI的机构,安全监管能力只剩一半
- Beyond Imitation: Filtering On-Policy Distillation by Reasoning ProgressAI模仿老师模型学习时,会误伤本来推理正确的步骤,新方法专门过滤掉这种误伤
- PersonalBench: Measuring the Authorship Gap in LLM Personalization让AI模仿某人的文风,结果发现它始终摆脱不了自己的腔调
METAL MEDIA 最新报道
图片来源: Badal Nyalang et al., arXiv:2608.18094, CC BY 4.0