NE-BERT: A Multilingual Language Model for Nine Northeast Indian Languages
A from-scratch language model for 9 minority languages of Northeast India beats big multilingual models on understanding them
A researcher affiliated with MWire Labs trained a language model called NE-BERT on about 8.3 million sentences covering 9 indigenous Northeast Indian languages plus Hindi and English. To handle extremely data-poor languages like Pnar, which has only 1,002 sentences, the team used weighted sampling and built a custom tokenizer, the tool that splits text into pieces a model can process. The resulting model predicted text far better than existing Indic-focused models IndicBERT-V2 and MuRIL, by 15.97x and 7.64x respectively on average perplexity, a measure of how well a model predicts text.
METAL MEDIA explanatory visual
A from-scratch language model for 9 minority languages of Northeast India beats big multilingual models on understanding them
- 01Northeast India has over 200 distinct languages, yet existing multilingual language models barely cover them
- 02The team gathered about 8.3 million sentences across Assamese, Garo, Khasi, Meitei, Mizo, Naga, Nyishi, Pnar, and Kokborok plus Hindi and English, and upsampled ultra-scarce languages like Pnar (1,002 sentences) and Kokborok (2,463 sentences) by 100x during tokenizer training to prevent their vocabulary from being fragmented into useless character-level pieces
- 03Instead of the common Byte-Pair Encoding method, they chose SentencePiece Unigram to build a custom 50,368-token tokenizer, which split words 1.50x more efficiently than mBERT's tokenizer
- 04On perplexity (lower is better), NE-BERT averaged 2.21 across the 9 Northeast Indian languages, far ahead of IndicBERT-V2's 35.29 and MuRIL's 16.88, with especially large gaps on ultra-low-resource languages like Pnar and Nyishi where the baseline models nearly broke down
- 05On a real downstream task, part-of-speech tagging, NE-BERT averaged 82.4% accuracy across Khasi, Mizo, and Nagamese, beating mBERT by 9.1 percentage points and IndicBERT-V2 by 23.2 points, while training the whole model cost just $7.31 on a single A40 GPU over about 17 hours
What they did
- Northeast India has over 200 distinct languages, yet existing multilingual language models barely cover them
- The team gathered about 8.3 million sentences across Assamese, Garo, Khasi, Meitei, Mizo, Naga, Nyishi, Pnar, and Kokborok plus Hindi and English, and upsampled ultra-scarce languages like Pnar (1,002 sentences) and Kokborok (2,463 sentences) by 100x during tokenizer training to prevent their vocabulary from being fragmented into useless character-level pieces
- Instead of the common Byte-Pair Encoding method, they chose SentencePiece Unigram to build a custom 50,368-token tokenizer, which split words 1.50x more efficiently than mBERT's tokenizer
- On perplexity (lower is better), NE-BERT averaged 2.21 across the 9 Northeast Indian languages, far ahead of IndicBERT-V2's 35.29 and MuRIL's 16.88, with especially large gaps on ultra-low-resource languages like Pnar and Nyishi where the baseline models nearly broke down
- On a real downstream task, part-of-speech tagging, NE-BERT averaged 82.4% accuracy across Khasi, Mizo, and Nagamese, beating mBERT by 9.1 percentage points and IndicBERT-V2 by 23.2 points, while training the whole model cost just $7.31 on a single A40 GPU over about 17 hours

| Language | ISO | Sentences | Tokens | Virtual Count | Weight | Family | Source |
|---|---|---|---|---|---|---|---|
| Anchor Languages | |||||||
| Hindi | hin | 3,404,007 | — | 170,200 | 0.05× | Indo-Aryan | HF Datasets |
| English | eng | 500,000 | — | 100,000 | 0.2× | Germanic | HF Datasets |
| Northeast Indian Languages | |||||||
| Meitei | mni | 1,354,323 | 42,504,181 | 1,354,323 | 1.0× | Sino-Tibetan | Curated |
| Assamese | asm | 1,000,000 | 38,652,391 | 1,000,000 | 1.0× | Indo-Aryan | Curated |
| Khasi | kha | 1,000,000 | 17,472,606 | 1,000,000 | 1.0× | Austroasiatic | Curated |
| Mizo | lus | 1,000,000 | 26,774,164 | 1,000,000 | 1.0× | Sino-Tibetan | Curated |
| Nyishi | njz | 55,870 | 560,374 | 1,117,400 | 20.0× | Sino-Tibetan | WMT 2025 |
| Naga | nag | 13,918 | 508,980 | 278,360 | 20.0× | Sino-Tibetan | Curated |
| Garo | grt | 10,817 | 243,251 | 216,340 | 20.0× | Sino-Tibetan | Curated |
| Kokborok | trp | 2,463 | 89,851 | 246,300 | 100.0× | Sino-Tibetan | WMT 2025 |
| Pnar | pbv | 1,002 | 52,144 | 100,200 | 100.0× | Austroasiatic | Curated |
| NE Total | 4,438,393 | 126,857,942 | |||||
| Overall Total | 8,342,400 | — | 6,583,123 |
| Model | Params | Layers | Hidden | Heads |
|---|---|---|---|---|
| NE-BERT | 149M | 22 | 768 | 12 |
| mBERT | 110M | 12 | 768 | 12 |
| IndicBERT-V2 | 237M | 12 | 1024 | 16 |
| MuRIL | 236M | 24 | 1024 | 16 |
| Language | NE-BERT | IB-V2 | MuRIL | mBERT |
|---|---|---|---|---|
| Assamese | 1.76 | 9.01 | 5.62 | 1.65 |
| Meitei | 1.89 | 3.77 | 3.22 | 1.44 |
| Khasi | 1.27 | 2.40 | 1.93 | 1.36 |
| Mizo | 1.90 | 8.95 | 7.00 | 2.44 |
| Garo | 2.64 | 26.37 | 15.88 | 3.64 |
| Kokborok | 1.72 | 9.15 | 5.41 | 2.23 |
| Pnar | 2.92 | 66.92 | 39.65 | 5.30 |
| Naga | 1.49 | 3.83 | 3.39 | 1.81 |
| Nyishi | 4.33 | 187.20 | 75.80 | 6.05 |
| English | 1.55 | 14.81 | 5.57 | 2.64 |
| Hindi | 1.43 | 10.08 | 5.75 | 1.79 |
| NE Avg. | 2.21 | 35.29 | 16.88 | 2.77 |
| Overall | 2.17 | 31.14 | 15.38 | 2.76 |
| Language | NE-B | IB-V2 | MuRIL | mB |
|---|---|---|---|---|
| Assamese | 1.63 | 1.61 | 1.72 | 3.80 |
| Meitei | 1.52 | 2.77 | 3.01 | 4.17 |
| Khasi | 1.31 | 1.77 | 1.78 | 1.78 |
| Mizo | 1.36 | 1.73 | 1.79 | 1.78 |
| Garo | 2.37 | 2.84 | 2.80 | 2.89 |
| Kokborok | 2.07 | 2.07 | 2.22 | 2.19 |
| Pnar | 1.61 | 1.69 | 1.69 | 1.67 |
| Naga | 1.56 | 1.93 | 1.77 | 1.97 |
| Nyishi | 1.67 | 2.34 | 2.44 | 2.36 |
| NE Avg. | 1.68 | 2.08 | 2.14 | 2.51 |
| Language | NE-B | IB-V2 | MuR | mB |
|---|---|---|---|---|
| Assamese | 0.230 | 0.880 | 0.731 | 0.442 |
| Meitei | 0.220 | 0.815 | 0.778 | 0.331 |
| Khasi | 0.092 | 0.445 | 0.336 | 0.158 |
| Mizo | 0.272 | 1.163 | 1.063 | 0.486 |
| Garo | 0.503 | 1.997 | 1.663 | 0.800 |
| Kokborok | 0.277 | 1.136 | 0.926 | 0.433 |
| Pnar | 0.680 | 2.795 | 2.446 | 1.092 |
| Naga | 0.171 | 0.709 | 0.594 | 0.321 |
| Nyishi | 0.789 | 3.763 | 3.231 | 1.306 |
| English | 0.190 | 0.862 | 0.558 | 0.328 |
| Hindi | 0.192 | 0.863 | 0.659 | 0.345 |
| NE Avg. | 0.359 | 1.545 | 1.307 | 0.597 |
| Overall | 0.347 | 1.497 | 1.271 | 0.590 |
| Language | NE-BERT | mBERT | IB-V2 | MuRIL |
|---|---|---|---|---|
| Khasi | 87.7 | 82.3 | 80.1 | 61.2 |
| Mizo | 73.2 | 66.3 | 55.7 | 43.1 |
| Nagamese | 86.4 | 71.3 | 41.7 | 44.8 |
| Average | 82.4 | 73.3 | 59.2 | 49.7 |
Why it matters
This work shows that for languages with almost no digital data, careful data balancing and tokenizer design can matter more than simply scaling up model size, making it possible to build effective language models cheaply. By releasing the model, training code, and datasets under CC-BY-4.0, it gives underserved language communities a concrete starting point to build their own NLP tools.
Terms in this paper
- Perplexity · a score measuring how well a language model predicts text; lower means better prediction
- Tokenization fertility · the average number of pieces (tokens) a tokenizer splits each word into; lower means more efficient
- SentencePiece Unigram · a method for splitting text into subword pieces using probabilities, which preserves complex word structures better than greedy methods
- Weighted sampling · artificially increasing the representation of data-scarce languages during training to fix imbalance
- Masked language modeling (MLM) · a training method where parts of a sentence are hidden and the model learns to guess the missing words
Original abstract (English)
Large pretrained language models have demonstrated remarkable capabilities across diverse languages, yet critically underrepresented low-resource languages remain marginalized. We present NE-BERT, a domain-specific multilingual encoder model trained on approximately 8.3 million sentences spanning 9 Northeast Indian languages and 2 anchor languages (Hindi, English), a linguistically diverse region with minimal representation in existing multilingual models. By employing weighted data sampling and a custom SentencePiece Unigram tokenizer, NE-BERT outperforms IndicBERT-V2 and MuRIL across all 9 Northeast Indian languages, achieving 15.97X and 7.64X lower average perplexity respectively, with 1.50X better tokenization fertility than mBERT. We address critical vocabulary fragmentation issues in extremely low-resource languages such as Pnar (1,002 sentences) and Kokborok (2,463 sentences) through aggressive upsampling strategies. Downstream evaluation on part-of-speech tagging validates practical utility on three Northeast Indian languages. We release NE-BERT, test sets, and training corpus under CC-BY-4.0 to support NLP research and digital inclusion for Northeast Indian communities.
Read on arXivLatest papers
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?AI coding agents were tested on fixing real scientific software, and even the best one failed more than half the time
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM ServingMaking sparse attention fast enough and accurate enough for real LLM serving, not just papers
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM AgentsMaking customer-service AI agents follow the whole procedure, not just avoid one bad action
- EXIMO: VLM Guided Exploration of VLA PoliciesTeaching a robot new chores without human teleoperation, by letting a chatty AI supervise it
- EnvHarness: Awakening Static Worlds for Agent LearningInstead of building new training worlds from scratch, this work adds a plug-in layer that reshapes existing ones around each agent's actual weaknesses
- Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the ModelCompanies that rent AI instead of owning it can only do half of AI safety oversight
- Beyond Imitation: Filtering On-Policy Distillation by Reasoning ProgressA fix for AI models that get penalized by their teacher even when they're reasoning correctly
- PersonalBench: Measuring the Authorship Gap in LLM PersonalizationAI can be prompted to write 'like someone,' but its own voice never fully disappears
Latest from METAL MEDIA
Figures: Badal Nyalang et al., arXiv:2608.18094, CC BY 4.0