A Speech Corpus for Mizo Automatic Speech Recognition: Whisper and SraVaani 1.0 Fine-Tuning with Morphology-Aware Evaluation
为印度低资源语言米佐语打造语音识别:用17.62小时新数据微调Whisper和SraVaani
研究团队从印度米佐拉姆邦200名说话人处收集并整理了17.62小时的米佐语语音数据,用这些数据分别微调了本不支持米佐语的Whisper模型和已经支持米佐语的印度多语言模型SraVaani 1.0。Whisper-large-v3取得了最低的常规词错误率18.08%,在专为米佐语设计的形态感知评估下进一步降到7.22%。SraVaani 1.0经过微调后,词错误率从58.27%大幅降至29.45%。
METAL MEDIA 解读图
为印度低资源语言米佐语打造语音识别:用17.62小时新数据微调Whisper和SraVaani
- 01200名说话人通过网页界面朗读了从报纸和法院判决书翻译来的约8000个米佐语句子,最终整理出8274条句子级录音,总计17.62小时。
- 02研究者分别对Whisper的small、medium、large-v3三种规模模型,以及本身已支持米佐语的印度多语言模型SraVaani 1.0进行微调,并用不重叠说话人的训练、验证、测试集进行评估。
- 03由于米佐语中形态边界的空格写法可自由变化(同一词素既可连写也可分写),常规词错误率会高估实际错误,研究者因此设计了允许最多合并四个相邻词进行比较的形态感知词错误率。
- 04Whisper-large-v3表现最佳,常规词错误率为18.08%,形态感知词错误率为7.22%;SraVaani 1.0零样本测试时词错误率高达58.27%,经米佐语专门微调后降至29.45%,形态感知词错误率为17.93%。
- 05错误分析显示,未经微调的SraVaani 1.0有时会输出完全错误的文字系统(如梅泰文或天城文而非米佐文字),并在人名地名识别、声门塞音识别上出现较多错误,微调后这些错误大幅减少。
他们做了什么
- 200名说话人通过网页界面朗读了从报纸和法院判决书翻译来的约8000个米佐语句子,最终整理出8274条句子级录音,总计17.62小时。
- 研究者分别对Whisper的small、medium、large-v3三种规模模型,以及本身已支持米佐语的印度多语言模型SraVaani 1.0进行微调,并用不重叠说话人的训练、验证、测试集进行评估。
- 由于米佐语中形态边界的空格写法可自由变化(同一词素既可连写也可分写),常规词错误率会高估实际错误,研究者因此设计了允许最多合并四个相邻词进行比较的形态感知词错误率。
- Whisper-large-v3表现最佳,常规词错误率为18.08%,形态感知词错误率为7.22%;SraVaani 1.0零样本测试时词错误率高达58.27%,经米佐语专门微调后降至29.45%,形态感知词错误率为17.93%。
- 错误分析显示,未经微调的SraVaani 1.0有时会输出完全错误的文字系统(如梅泰文或天城文而非米佐文字),并在人名地名识别、声门塞音识别上出现较多错误,微调后这些错误大幅减少。

| Speakers | Sentences | Hours | |
|---|---|---|---|
| Training | 184 | 7656 | 16.18 |
| Validation | 11 | 426 | 1.02 |
| Testing | 05 | 192 | 0.42 |

| Sentences | 8274 |
|---|---|
| Total duration | 17.62 hours |
| Minimum duration | 0.63 seconds |
| Maximum duration | 41.22 seconds |
| Mean duration | 7.67 seconds |
| Median duration | 6.94 seconds |
| Model | Architecture | Parameters | Pretraining data | Mel bins |
|---|---|---|---|---|
| Whisper-small | Encoder–Decoder Transformer | 244 M | 680k h | 80 |
| Whisper-medium | Encoder–Decoder Transformer | 769 M | 680k h | 80 |
| Whisper-large-v3 | Encoder–Decoder Transformer | 1,550 M | ∼5M h | 128 |
| SraVaani 1.0 | FastConformer Hybrid RNNT/CTC | ∼430 M | ∼31k h | – |
| Model | Optimizer | LR | Effective batch size | Epochs | Selection |
|---|---|---|---|---|---|
| Whisper-small | Adafactor | 5×10−6 | 16 | 20 | Best val. WER |
| Whisper-medium | Adafactor | 5×10−6 | 16 | 20 | Best val. WER |
| Whisper-large-v3 | Adafactor | 5×10−6 | 16 | 20 | Best val. WER |
| SraVaani 1.0 | AdamW | 1×10−4 | 16 | 20 + extended | Best val. WER |
| Model | Best epoch | Validation WER |
|---|---|---|
| Whisper-small Mizo–FT | 15 | 28.99 |
| Whisper-medium Mizo–FT | 13 | 26.51 |
| Whisper-large-v3 Mizo–FT | 13 | 23.00 |
| SraVaani 1.0 Mizo–FT | 18 | 33.81 |
| Model | CER (%) | WER (%) | MA-WER (%) |
|---|---|---|---|
| Whisper-small Mizo–FT | 04.83 | 24.00 | 11.49 |
| Whisper-medium Mizo–FT | 04.02 | 21.69 | 08.87 |
| Whisper-large-v3 Mizo–FT | 03.26 | 18.08 | 07.22 |
| SraVaani 1.0 | 17.71 | 58.27 | 36.27 |
| SraVaani 1.0 Mizo-FT | 06.90 | 29.45 | 17.93 |
| Model | Foreign script | Names | Glottal stops | Numeral transcripts | Code-mix error | < t Ω > |
|---|---|---|---|---|---|---|
| Whisper-small Mizo–FT | NIL | 17 | 7 | 2 | 9 | 0 |
| Whisper-medium Mizo–FT | NIL | 18 | 6 | 5 | 3 | 0 |
| Whisper-large-v3 Mizo–FT | 4 sentences | 8 | 4 | 3 | 2 | 0 |
| SraVaani 1.0 | 21 sentences | 44 | 9 | 19 | 49 | 31 |
| SraVaani 1.0 Mizo–FT | NIL | 24 | 6 | 2 | 19 | 4 |
为什么重要
这项工作公开发布了语音语料库、微调后的语音识别模型,以及一套适合米佐语特点的评估指标,为米佐语语音技术的后续开发提供了基础。研究还表明,对于形态边界空格写法灵活的语言,常规词错误率可能高估真实识别质量,这对使用罗马字母的其他藏缅语族语言评估也有参考价值。
本文术语
- 词错误率(WER) · 通过比较识别结果与参考文本中的替换、删除、插入词数来衡量语音识别准确度的标准指标
- 形态感知词错误率(MA-WER) · 研究者为容忍米佐语形态边界空格写法差异而设计的新指标,允许最多合并四个相邻词后再比较
- 字符错误率(CER) · 与词错误率计算方式相同,但以字符而非单词为基本比较单位
- 零样本评估 · 不对某语言进行额外微调,直接用预训练模型测试其表现的方式
- Whisper / SraVaani 1.0 · Whisper是基于Transformer的多语言语音识别与翻译模型;SraVaani 1.0是覆盖多种印度语言的多语言语音识别模型
论文原文摘要(英文)
This study reports the development of an Automatic Speech Recognition (ASR) system in Mizo, a low-resource language. The development included collecting 17.62 hours of speech data, curating it, and fine-tuning the Mizo ASR system with three Whisper multilingual models and with the SraVaani 1.0 Indic multilingual model. Whisper-large-v3 achieved the lowest conventional WER (18.08%), while morphology-aware evaluation yielded a WER of 7.22%. Zero-shot evaluation of the SraVaani 1.0 Indic multilingual model yielded a WER of 58.27%, while Mizo-specific fine-tuning reduced the conventional WER to 29.45% and the morphology-aware WER to 17.93%. The results demonstrate that the Whisper model can achieve a substantially low WER, even when adapted to an unseen language. In contrast, SraVaani 1.0 supports the Mizo language in its multilingual model; however, fine-tuning with carefully curated Mizo speech data substantially improves its performance.
在 arXiv 阅读最新论文
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?让AI编程助手去修复真实科学软件,连最强的那个也有一半以上任务没做对
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving把稀疏注意力从论文原型变成能真正上线服务的加速方案
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents让客服AI坐席不只是拦住一个危险动作,而是把整个流程走对
- EXIMO: VLM Guided Exploration of VLA Policies不用人工遥控演示,让会说话的AI来教机械臂做新家务
- EnvHarness: Awakening Static Worlds for Agent Learning不重新搭建训练环境,而是给现有环境套一层可插拔组件,针对每个智能体的具体弱点重新塑形
- Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the Model租用AI而非拥有AI的机构,安全监管能力只剩一半
- Beyond Imitation: Filtering On-Policy Distillation by Reasoning ProgressAI模仿老师模型学习时,会误伤本来推理正确的步骤,新方法专门过滤掉这种误伤
- PersonalBench: Measuring the Authorship Gap in LLM Personalization让AI模仿某人的文风,结果发现它始终摆脱不了自己的腔调
METAL MEDIA 最新报道
图片来源: Priyankoo Sarmah et al., arXiv:2608.19361, cc-by-nc-nd-4.0