인도어 토크나이저가 단어 뿌리를 함부로 자르는 문제를 형태소 경계를 지키는 방식으로 고친 연구
arXiv:2608.180872026-08-20
SuTRA : Structurally-Unified Tokenization with Root Awareness
인도어 토크나이저가 단어 뿌리를 함부로 자르는 문제를 형태소 경계를 지키는 방식으로 고친 연구
기존 BPE 같은 서브워드 토크나이저는 통계적 압축에만 최적화되어 있어서 힌디어, 마라티어, 구자라트어 같은 인도어에서 단어의 어근과 접사를 제멋대로 잘라버린다. 연구팀은 이런 현상을 형태소 파쇄라고 이름 붙이고, 문자 결합 규칙과 형태소 경계 정보를 함께 반영하는 SuTRA라는 토크나이저를 제안했다. 새로 만든 대규모 정답 데이터셋으로 검증한 결과 경계 정확도, 의미 복원력, 기계번역 성능이 모두 기존 BPE보다 개선되었다.
METAL MEDIA 해설 도표
인도어 토크나이저가 단어 뿌리를 함부로 자르는 문제를 형태소 경계를 지키는 방식으로 고친 연구
01기존 BPE, WordPiece 같은 서브워드 토크나이저는 빈도만 보고 단어를 쪼개서 힌디어 asādhāraṇ 같은 단어에서 부정 접두사 a가 어근 sādhāraṇ에 붙어버리는 형태소 파쇄가 발생한다
02SuTRA는 두 단계로 작동한다. 1단계는 데바나가리 같은 문자 체계에서 자음과 모음 기호가 합쳐진 음절 단위(악샤라)가 쪼개지지 않도록 규칙을 적용하고, 사전이나 seq2seq 모델로 형태소 경계를 표시한다. 2단계는 BPE 방식의 병합 점수에 형태소 경계를 침범하는 병합을 벌점 주는 항을 넣고, 학습이 진행될수록 이 제약을 점점 완화한다
03힌디어, 마라티어, 구자라트어 약 56만 단어 규모의 정답 형태소 분할 데이터셋을 직접 구축했다
04경계 F1 점수에서 BPE 대비 최대 14.7퍼센트포인트, 의미 복원력에서 힌디어 기준 34퍼센트 향상되었고 기계번역에서는 평균 8.08 chrF2 점수가 올랐다
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.
무엇을 했나
기존 BPE, WordPiece 같은 서브워드 토크나이저는 빈도만 보고 단어를 쪼개서 힌디어 asādhāraṇ 같은 단어에서 부정 접두사 a가 어근 sādhāraṇ에 붙어버리는 형태소 파쇄가 발생한다
SuTRA는 두 단계로 작동한다. 1단계는 데바나가리 같은 문자 체계에서 자음과 모음 기호가 합쳐진 음절 단위(악샤라)가 쪼개지지 않도록 규칙을 적용하고, 사전이나 seq2seq 모델로 형태소 경계를 표시한다. 2단계는 BPE 방식의 병합 점수에 형태소 경계를 침범하는 병합을 벌점 주는 항을 넣고, 학습이 진행될수록 이 제약을 점점 완화한다
힌디어, 마라티어, 구자라트어 약 56만 단어 규모의 정답 형태소 분할 데이터셋을 직접 구축했다
경계 F1 점수에서 BPE 대비 최대 14.7퍼센트포인트, 의미 복원력에서 힌디어 기준 34퍼센트 향상되었고 기계번역에서는 평균 8.08 chrF2 점수가 올랐다
Figure 1: Morphological Shattering vs. Root Preservation. For the Hindi word asādhāraṇ, standard Bpe fuses the negation prefix into the root ([asā]+[dhāraṇ]), while SuTRA cleanly separates prefix and root ([a]+[sādhāraṇ]). We term such frequency-driven prefix–root fusion Morphological Shattering; SuTRA's root-preserving segmentations yield more stable subword units and reduce semantic blindness. Figure generated using PaperBanana [paperbanana].
Table 1: Morphological Datasets. We provide the first high-scale, LLM-verified morphological coverage for Indic scripts.
Figure 2: Overview of SuTRA. Phase 1 (Pre-tokenization) applies orthographic rules Φ to map each word into akshara-like units and uses a gold morphological lexicon or a seq2seq model to mark forbidden boundaries (morpheme boundaries that merges should not cross). Phase 2 (Morphology-Aware Merging) runs a BPE-style algorithm with scores S(a,b)=f(a,b)Ψ(a,b)γt, where Ψ downweights merges that violate forbidden boundaries and γt controls rigidity over training, biasing vocabulary toward merges that respect Indic script structure and morpheme boundaries. Figure generated using PaperBanana [paperbanana].
Table 2: Gold Standard Statistics. Hybrid of rule-based extraction and LLM-verified morphological segmentation.
Language
Source
Unique Words
Hindi
IndicCorp [ai4bharat_corpus]
≃ 160,000
Marathi
IndicCorp [ai4bharat_corpus]
≃ 200,000
Gujarati
IndicCorp [ai4bharat_corpus]
≃ 200,000
Total
≃ 560,000
Figure 4: Sample examples from the dataset used for Hindi to Marathi Machine Translation Task
Table 3: Morphological Alignment Evaluation. Evaluation of Boundary F1 and Fertility Ratio across Hindi, Marathi, and Gujarati. Higher F1 indicates superior structural integrity, while controlled fertility prevents arbitrary character-level fragmentation. Best F1 scores are in bold.
Tokenizer
Hindi
Marathi
Gujarati
F1 ↑
Fert. ↓
F1 ↑
Fert. ↓
F1 ↑
Fert. ↓
BPE (ACL'16)
0.482
1.285
0.470
1.225
0.591
1.126
WordPiece (arXiv'12)
0.411
1.214
0.527
1.300
0.596
1.173
SentencePiece (EMNLP'18)
0.438
1.315
0.083
1.183
0.591
1.156
Unigram (ACL'18)
0.439
1.310
0.507
1.256
0.669
1.137
SuperBPE (COLM'25)
0.089
2.517
0.084
3.084
0.096
3.113
MorphTok (ICML-W'25)
0.190
2.601
0.216
3.482
0.247
3.494
\rowcolor[gray]0.9 SuTRA (Ours)
0.586
1.412
0.617
1.755
0.584
1.454
Figure 5: Loss curves for BPE and SuTRA
Table 4: Semantic Recoverability (R2). Linear models evaluate immediate compositionality, while deeper MLP models test the preservation of recoverable structural signals.
Language
Linear R2 (Layer 0)
MLP R2 (Layer 2+)
SuTRA
Bpe
SuTRA
Bpe
Hindi
0.4464
0.3329
0.5048
0.3358
Marathi
0.4634
0.4619
0.5331
0.4649
Gujarati
0.4624
0.4640
0.5055
0.4510
Table 5: Machine Translation. SuTRA achieves the highest scores in Marathi → Hindi and remains comparable to the strongest baseline in the reverse task, demonstrating that root contiguity improves cross-lingual alignment.
Tokenizer
Hindi → Marathi
Marathi → Hindi
chrF2 ↑
COMET ↑
chrF2 ↑
COMET ↑
\rowcolor[gray]0.95 Statistical Baselines
Bpe (ACL'16)
24.62
0.5054
36.55
0.6253
WordPiece (arXiv'12)
30.37
0.6093
27.18
0.5223
SentencePiece (EMNLP'18)
25.01
0.4964
28.10
0.5147
Unigram (ACL'18)
26.55
0.5157
29.19
0.5349
\rowcolor[gray]0.95 Morphological Baselines
SuperBpe (COLM'25)
15.84
0.3047
16.27
0.3422
MorphTok (ICML-W'25)
26.75
0.5750
29.42
0.6007
\rowcolor[gray]0.9 SuTRA (Ours)
29.96
0.5752
38.84
0.6554
Table 6: Morphological Robustness Comparison. Results show that SuTRA effectively mitigates root fragmentation.
Tokenizer
Hindi (HI)
Marathi (MR)
Gujarati (GU)
Jac.(↑)
R.Aff.(↓)
Jac.(↑)
R.Aff.(↓)
Jac.(↑)
R.Aff.(↓)
Bpe
0.386
0.228
0.305
0.324
0.319
0.296
Unigram
0.363
0.232
0.294
0.340
0.320
0.308
WordPiece
0.413
0.249
0.293
0.355
0.314
0.316
SentencePiece
0.371
0.235
0.298
0.325
0.321
0.287
SuperBpe
0.624
0.146
0.655
0.127
0.631
0.126
MorphTok
0.801
0.067
0.823
0.055
0.792
0.051
\rowcolor[gray]0.9 SuTRA (Ours)
0.875
0.042
0.885
0.038
0.788
0.040
Table 7: Effect of the Annealed Rigidity Constraint. At the start of training (t=0,γ=4), the exponential penalty prioritizes structurally safe merges (Ψ≈1.0), forcing the tokenizer to build valid semantic roots despite lower raw frequencies. By the end of training (t=T,γ=0), the penalty disappears (Ψ0=1), reducing the score to standard Bpe frequency to optimize compression.
Candidate Merge (a,b)
Freq. f
Validity Ψ
Effective Score S(a,b)
t=0 (γ=4)
t=T (γ=0)
Safe Merge (e.g., intra-root)
1000
1.00
𝟏𝟎𝟎𝟎
1000
Boundary Violation (e.g., prefix+root)
1200
0.80
491
𝟏𝟐𝟎𝟎
Severe Violation (e.g., shattered prefix)
1500
0.50
93
𝟏𝟓𝟎𝟎
Table 8: Time Complexity Summary. Comparison of training and inference time complexity. Variables: N (corpus size), V (vocab size), M (unique candidate pairs), Vin (unique corpus words), and |w| (word length).
Tokenizer
Training
Inference (per word)
Standard Bpe
𝒪(N+VlogM)
𝒪(|w|)
MorphTok
𝒪(Vin⋅|w|2+N+VlogM)
𝒪(|w|2)
SuTRA (Ours)
𝒪(Vin⋅|w|2+N+VlogM)
𝒪(|w|)
Table 9: Comparison of Tokenizers (Normalized) across Hindi, Marathi, and Gujarati. EM: Exact Match Acc, Prec: Boundary Precision, Rec: Boundary Recall, F1: Boundary F1, Fert: Fertility Ratio. Best F1 scores are in bold.
Tokenizer
Hindi
Marathi
Gujarati
EM
Prec.
Rec.
F1
Fert.
EM
Prec.
Rec.
F1
Fert.
EM
Prec.
Rec.
F1
Fert.
BPE
0.271
0.398
0.609
0.482
1.285
0.268
0.396
0.578
0.470
1.225
0.396
0.534
0.662
0.591
1.126
SentencePiece
0.191
0.357
0.566
0.438
1.315
0.000
0.043
1.000
0.083
1.183
0.378
0.523
0.679
0.591
1.156
SuperBPE
0.000
0.056
0.215
0.089
2.517
0.000
0.050
0.264
0.084
3.084
0.000
0.058
0.289
0.096
3.113
Unigram
0.185
0.359
0.566
0.439
1.310
0.394
0.511
0.778
0.507
1.256
0.460
0.600
0.756
0.669
1.137
WordPiece
0.166
0.353
0.493
0.411
1.214
0.274
0.427
0.688
0.527
1.300
0.378
0.522
0.694
0.596
1.173
MorphTok
0.004
0.122
0.531
0.190
2.601
0.008
0.126
0.763
0.216
3.482
0.006
0.145
0.835
0.247
3.494
SuTRA
0.212
0.459
0.810
0.586
1.412
0.132
0.354
0.898
0.617
1.755
0.213
0.448
0.836
0.584
1.454
Table 10: Hyperparameters for the downstream Causal Language Modeling task.
Hyperparameter
Value
Architecture
GPT-2 (124M parameters)
Number of Layers
12
Number of Heads
12
Embedding Dimension (dmodel)
768
Context Window (nctx)
1024
Vocabulary Size
Language-dependent
Optimizer
AdamW
Learning Rate
5×10−4
Weight Decay
0.01
Warmup Steps
1,000
Target Training Tokens
1.5 Billion
Batch Size (per device)
16
Precision
Mixed (bf16 or fp16)
Table 11: Downstream performance on Causal Language Modeling (CLM). We report average Negative Log-Likelihood (Avg NLL) and Perplexity (PPL) on out-of-distribution (OOD) corpora. Lower values indicate better compression and predictive efficiency.
Language
Tokenizer
Vocab Size
Avg NLL ↓
Perplexity ↓
Hindi (HI)
BPE
32000
3.116
18.87
SuTRA
32000
3.016
16.08
Table 12: Morphological Shattering Comparison. Results show that SuTRA effectively mitigates root fragmentation.
Tokenizer
Hindi (HI)
Marathi (MR)
Gujarati (GU)
Jac.(↑)
R.Aff.(↓)
Jac.(↑)
R.Aff.(↓)
Jac.(↑)
R.Aff.(↓)
Bpe
0.386
0.228
0.305
0.324
0.319
0.296
Unigram
0.363
0.232
0.294
0.340
0.320
0.308
WordPiece
0.413
0.249
0.293
0.355
0.314
0.316
SentencePiece
0.371
0.235
0.298
0.325
0.321
0.287
SuperBpe
0.624
0.146
0.655
0.127
0.631
0.126
MorphTok
0.801
0.067
0.823
0.055
0.792
0.051
\rowcolor[gray]0.9 SuTRA (Ours)
0.875
0.042
0.885
0.038
0.788
0.040
왜 중요한가
토크나이저는 언어모델이 텍스트를 처리하는 첫 관문인데, 이 관문에서 단어의 의미 단위가 깨지면 모델이 아무리 커도 기본적인 의미 파악에 낭비되는 계산이 생긴다. 형태가 복잡한 언어를 다루는 다국어 모델을 만들 때 이 연구의 데이터셋과 방법론이 실질적인 참고 자료가 될 수 있다.
이 논문의 용어
서브워드 토크나이저 · 단어를 통계적으로 더 작은 조각으로 나눠 어휘 목록을 만드는 도구, 대표적으로 BPE가 있음
악샤라 · 데바나가리 등 인도 문자에서 자음과 모음 기호가 합쳐져 하나의 시각적 단위를 이루는 음절
형태소 경계 · 어근과 접사처럼 의미 단위가 나뉘는 지점
산디 · 인도아리아어군에서 단어나 형태소 경계의 소리가 합쳐지거나 변하는 현상
경계 F1 · 토크나이저가 예측한 분할 지점이 정답 형태소 경계와 얼마나 일치하는지 정밀도와 재현율을 함께 계산한 점수
본문에 싣지 못한 그림
Figure 3: Qualitative Comparison of Morphological Segmentation Across Tokenizers. SuTRA consistently matches the Gold Standard by respecting phonetic and morphological boundaries.