Everything K-culture — comebacks to K-beauty, straight to your inboxGet it in your inbox›
SuTRA : Structurally-Unified Tokenization with Root Awareness
arXiv:2608.180872026-08-20
A tokenizer that stops shredding word roots in Indic languages by respecting morpheme boundaries
Standard subword tokenizers like BPE optimize purely for statistical compression, which causes them to arbitrarily split roots and affixes in morphologically rich Indic languages such as Hindi, Marathi, and Gujarati, a problem the authors call Morphological Shattering. They propose SuTRA, which keeps script-level syllable units intact and penalizes merges that cross morpheme boundaries during BPE-style training. Tested against a new large-scale gold-standard dataset the team built, SuTRA improves boundary alignment, semantic recoverability, and machine translation quality over standard BPE.
METAL MEDIA explanatory visual
A tokenizer that stops shredding word roots in Indic languages by respecting morpheme boundaries
01Standard tokenizers like BPE and WordPiece split words based purely on frequency, causing errors like fusing the negation prefix 'a' into the root in the Hindi word asādhāraṇ instead of keeping them separate
02SuTRA works in two phases: Phase 1 applies orthographic rules so that akshara-like syllable units (consonant plus vowel mark combinations found in scripts like Devanagari) are not broken apart, and marks forbidden boundaries using a lexicon or a seq2seq model for unknown words; Phase 2 runs BPE-style merging but penalizes merges that cross those forbidden boundaries, with the penalty strength gradually relaxed over training
03The team built a new gold-standard morphological segmentation dataset of about 560,000 verified words across Hindi, Marathi, and Gujarati using a hybrid rule-based and LLM-verified pipeline
04SuTRA achieves up to +14.7 percentage points in Boundary F1 and +34% in semantic recoverability for Hindi compared to BPE, and an average +8.08 chrF2 gain in machine translation
An explanatory diagram made by METAL MEDIA, not a figure supplied by the paper's authors.
What they did
Standard tokenizers like BPE and WordPiece split words based purely on frequency, causing errors like fusing the negation prefix 'a' into the root in the Hindi word asādhāraṇ instead of keeping them separate
SuTRA works in two phases: Phase 1 applies orthographic rules so that akshara-like syllable units (consonant plus vowel mark combinations found in scripts like Devanagari) are not broken apart, and marks forbidden boundaries using a lexicon or a seq2seq model for unknown words; Phase 2 runs BPE-style merging but penalizes merges that cross those forbidden boundaries, with the penalty strength gradually relaxed over training
The team built a new gold-standard morphological segmentation dataset of about 560,000 verified words across Hindi, Marathi, and Gujarati using a hybrid rule-based and LLM-verified pipeline
SuTRA achieves up to +14.7 percentage points in Boundary F1 and +34% in semantic recoverability for Hindi compared to BPE, and an average +8.08 chrF2 gain in machine translation
Figure 1: Morphological Shattering vs. Root Preservation. For the Hindi word asādhāraṇ, standard Bpe fuses the negation prefix into the root ([asā]+[dhāraṇ]), while SuTRA cleanly separates prefix and root ([a]+[sādhāraṇ]). We term such frequency-driven prefix–root fusion Morphological Shattering; SuTRA's root-preserving segmentations yield more stable subword units and reduce semantic blindness. Figure generated using PaperBanana [paperbanana].
Table 1: Morphological Datasets. We provide the first high-scale, LLM-verified morphological coverage for Indic scripts.
Figure 2: Overview of SuTRA. Phase 1 (Pre-tokenization) applies orthographic rules Φ to map each word into akshara-like units and uses a gold morphological lexicon or a seq2seq model to mark forbidden boundaries (morpheme boundaries that merges should not cross). Phase 2 (Morphology-Aware Merging) runs a BPE-style algorithm with scores S(a,b)=f(a,b)Ψ(a,b)γt, where Ψ downweights merges that violate forbidden boundaries and γt controls rigidity over training, biasing vocabulary toward merges that respect Indic script structure and morpheme boundaries. Figure generated using PaperBanana [paperbanana].
Table 2: Gold Standard Statistics. Hybrid of rule-based extraction and LLM-verified morphological segmentation.
Language
Source
Unique Words
Hindi
IndicCorp [ai4bharat_corpus]
≃ 160,000
Marathi
IndicCorp [ai4bharat_corpus]
≃ 200,000
Gujarati
IndicCorp [ai4bharat_corpus]
≃ 200,000
Total
≃ 560,000
Figure 4: Sample examples from the dataset used for Hindi to Marathi Machine Translation Task
Table 3: Morphological Alignment Evaluation. Evaluation of Boundary F1 and Fertility Ratio across Hindi, Marathi, and Gujarati. Higher F1 indicates superior structural integrity, while controlled fertility prevents arbitrary character-level fragmentation. Best F1 scores are in bold.
Tokenizer
Hindi
Marathi
Gujarati
F1 ↑
Fert. ↓
F1 ↑
Fert. ↓
F1 ↑
Fert. ↓
BPE (ACL'16)
0.482
1.285
0.470
1.225
0.591
1.126
WordPiece (arXiv'12)
0.411
1.214
0.527
1.300
0.596
1.173
SentencePiece (EMNLP'18)
0.438
1.315
0.083
1.183
0.591
1.156
Unigram (ACL'18)
0.439
1.310
0.507
1.256
0.669
1.137
SuperBPE (COLM'25)
0.089
2.517
0.084
3.084
0.096
3.113
MorphTok (ICML-W'25)
0.190
2.601
0.216
3.482
0.247
3.494
\rowcolor[gray]0.9 SuTRA (Ours)
0.586
1.412
0.617
1.755
0.584
1.454
Figure 5: Loss curves for BPE and SuTRA
Table 4: Semantic Recoverability (R2). Linear models evaluate immediate compositionality, while deeper MLP models test the preservation of recoverable structural signals.
Language
Linear R2 (Layer 0)
MLP R2 (Layer 2+)
SuTRA
Bpe
SuTRA
Bpe
Hindi
0.4464
0.3329
0.5048
0.3358
Marathi
0.4634
0.4619
0.5331
0.4649
Gujarati
0.4624
0.4640
0.5055
0.4510
Table 5: Machine Translation. SuTRA achieves the highest scores in Marathi → Hindi and remains comparable to the strongest baseline in the reverse task, demonstrating that root contiguity improves cross-lingual alignment.
Tokenizer
Hindi → Marathi
Marathi → Hindi
chrF2 ↑
COMET ↑
chrF2 ↑
COMET ↑
\rowcolor[gray]0.95 Statistical Baselines
Bpe (ACL'16)
24.62
0.5054
36.55
0.6253
WordPiece (arXiv'12)
30.37
0.6093
27.18
0.5223
SentencePiece (EMNLP'18)
25.01
0.4964
28.10
0.5147
Unigram (ACL'18)
26.55
0.5157
29.19
0.5349
\rowcolor[gray]0.95 Morphological Baselines
SuperBpe (COLM'25)
15.84
0.3047
16.27
0.3422
MorphTok (ICML-W'25)
26.75
0.5750
29.42
0.6007
\rowcolor[gray]0.9 SuTRA (Ours)
29.96
0.5752
38.84
0.6554
Table 6: Morphological Robustness Comparison. Results show that SuTRA effectively mitigates root fragmentation.
Tokenizer
Hindi (HI)
Marathi (MR)
Gujarati (GU)
Jac.(↑)
R.Aff.(↓)
Jac.(↑)
R.Aff.(↓)
Jac.(↑)
R.Aff.(↓)
Bpe
0.386
0.228
0.305
0.324
0.319
0.296
Unigram
0.363
0.232
0.294
0.340
0.320
0.308
WordPiece
0.413
0.249
0.293
0.355
0.314
0.316
SentencePiece
0.371
0.235
0.298
0.325
0.321
0.287
SuperBpe
0.624
0.146
0.655
0.127
0.631
0.126
MorphTok
0.801
0.067
0.823
0.055
0.792
0.051
\rowcolor[gray]0.9 SuTRA (Ours)
0.875
0.042
0.885
0.038
0.788
0.040
Table 7: Effect of the Annealed Rigidity Constraint. At the start of training (t=0,γ=4), the exponential penalty prioritizes structurally safe merges (Ψ≈1.0), forcing the tokenizer to build valid semantic roots despite lower raw frequencies. By the end of training (t=T,γ=0), the penalty disappears (Ψ0=1), reducing the score to standard Bpe frequency to optimize compression.
Candidate Merge (a,b)
Freq. f
Validity Ψ
Effective Score S(a,b)
t=0 (γ=4)
t=T (γ=0)
Safe Merge (e.g., intra-root)
1000
1.00
𝟏𝟎𝟎𝟎
1000
Boundary Violation (e.g., prefix+root)
1200
0.80
491
𝟏𝟐𝟎𝟎
Severe Violation (e.g., shattered prefix)
1500
0.50
93
𝟏𝟓𝟎𝟎
Table 8: Time Complexity Summary. Comparison of training and inference time complexity. Variables: N (corpus size), V (vocab size), M (unique candidate pairs), Vin (unique corpus words), and |w| (word length).
Tokenizer
Training
Inference (per word)
Standard Bpe
𝒪(N+VlogM)
𝒪(|w|)
MorphTok
𝒪(Vin⋅|w|2+N+VlogM)
𝒪(|w|2)
SuTRA (Ours)
𝒪(Vin⋅|w|2+N+VlogM)
𝒪(|w|)
Table 9: Comparison of Tokenizers (Normalized) across Hindi, Marathi, and Gujarati. EM: Exact Match Acc, Prec: Boundary Precision, Rec: Boundary Recall, F1: Boundary F1, Fert: Fertility Ratio. Best F1 scores are in bold.
Tokenizer
Hindi
Marathi
Gujarati
EM
Prec.
Rec.
F1
Fert.
EM
Prec.
Rec.
F1
Fert.
EM
Prec.
Rec.
F1
Fert.
BPE
0.271
0.398
0.609
0.482
1.285
0.268
0.396
0.578
0.470
1.225
0.396
0.534
0.662
0.591
1.126
SentencePiece
0.191
0.357
0.566
0.438
1.315
0.000
0.043
1.000
0.083
1.183
0.378
0.523
0.679
0.591
1.156
SuperBPE
0.000
0.056
0.215
0.089
2.517
0.000
0.050
0.264
0.084
3.084
0.000
0.058
0.289
0.096
3.113
Unigram
0.185
0.359
0.566
0.439
1.310
0.394
0.511
0.778
0.507
1.256
0.460
0.600
0.756
0.669
1.137
WordPiece
0.166
0.353
0.493
0.411
1.214
0.274
0.427
0.688
0.527
1.300
0.378
0.522
0.694
0.596
1.173
MorphTok
0.004
0.122
0.531
0.190
2.601
0.008
0.126
0.763
0.216
3.482
0.006
0.145
0.835
0.247
3.494
SuTRA
0.212
0.459
0.810
0.586
1.412
0.132
0.354
0.898
0.617
1.755
0.213
0.448
0.836
0.584
1.454
Table 10: Hyperparameters for the downstream Causal Language Modeling task.
Hyperparameter
Value
Architecture
GPT-2 (124M parameters)
Number of Layers
12
Number of Heads
12
Embedding Dimension (dmodel)
768
Context Window (nctx)
1024
Vocabulary Size
Language-dependent
Optimizer
AdamW
Learning Rate
5×10−4
Weight Decay
0.01
Warmup Steps
1,000
Target Training Tokens
1.5 Billion
Batch Size (per device)
16
Precision
Mixed (bf16 or fp16)
Table 11: Downstream performance on Causal Language Modeling (CLM). We report average Negative Log-Likelihood (Avg NLL) and Perplexity (PPL) on out-of-distribution (OOD) corpora. Lower values indicate better compression and predictive efficiency.
Language
Tokenizer
Vocab Size
Avg NLL ↓
Perplexity ↓
Hindi (HI)
BPE
32000
3.116
18.87
SuTRA
32000
3.016
16.08
Table 12: Morphological Shattering Comparison. Results show that SuTRA effectively mitigates root fragmentation.
Tokenizer
Hindi (HI)
Marathi (MR)
Gujarati (GU)
Jac.(↑)
R.Aff.(↓)
Jac.(↑)
R.Aff.(↓)
Jac.(↑)
R.Aff.(↓)
Bpe
0.386
0.228
0.305
0.324
0.319
0.296
Unigram
0.363
0.232
0.294
0.340
0.320
0.308
WordPiece
0.413
0.249
0.293
0.355
0.314
0.316
SentencePiece
0.371
0.235
0.298
0.325
0.321
0.287
SuperBpe
0.624
0.146
0.655
0.127
0.631
0.126
MorphTok
0.801
0.067
0.823
0.055
0.792
0.051
\rowcolor[gray]0.9 SuTRA (Ours)
0.875
0.042
0.885
0.038
0.788
0.040
Why it matters
Tokenizers are the first gate through which text enters a language model, so when that gate breaks meaningful word units apart, models waste capacity reconstructing basic semantics even if they are otherwise large and capable. This work's dataset and method offer a concrete reference point for anyone building multilingual models for morphologically complex languages.
Terms in this paper
subword tokenizer · a tool that splits words into smaller statistically frequent pieces to build a vocabulary, with BPE being a common example
akshara · an orthographic syllable unit in scripts like Devanagari formed by combining a consonant with a dependent vowel mark
morpheme boundary · the point where meaningful units like a root and an affix are divided
Sandhi · a phenomenon in Indo-Aryan languages where sounds at word or morpheme boundaries fuse or change
Boundary F1 · a score combining precision and recall of how well predicted split points match gold morpheme boundaries
Figures we cannot republish
Figure 3: Qualitative Comparison of Morphological Segmentation Across Tokenizers. SuTRA consistently matches the Gold Standard by respecting phonetic and morphological boundaries.
Existing subword tokenizers optimize statistical compression but ignore morphological structure, particularly the relationship between roots and affixes. This is harmful for morphologically rich Indic languages, where basic units are complex orthographic syllables (aksharas) rather than letters. Frequency-based methods over-fragment words, arbitrarily splitting roots and affixes - a phenomenon we term Morphological Shattering. We propose SuTRA (Structurally-Unified Tokenization with Root Awareness), a morphology-aware algorithm that preserves akshara indivisibility and penalizes merges crossing morphological boundaries. We also release a new morphological segmentation dataset for Hindi, Marathi, and Gujarati. SuTRA reduces shattering, achieving peak gains of +14.7% in morphological alignment (Boundary F1) and +34% in semantic recoverability (Hindi) over BPE. These structural gains yield an average improvement of +8.08 chrF2 in machine translation.