Language Models for Portuguese: A Systematic Mapping Study
首次系统梳理:面向葡萄牙语的46个AI语言模型全景图
巴西坎皮纳斯州立大学(UNICAMP)团队对2020年至2025年8月间发布的46个葡萄牙语语言模型进行了系统性调研。研究整理了每个模型的基础模型、架构、训练数据、许可协议以及代码、数据、权重的公开情况,并通过系统发育树的方式分析了模型之间的演化和衍生关系。研究还发现该领域存在明显空白,比如很少有论文讨论伦理考量或研究局限性。
METAL MEDIA 解读图
首次系统梳理:面向葡萄牙语的46个AI语言模型全景图
- 01采用系统性文献梳理方法,结合滚雪球式检索(顺着已知论文的引用关系向前向后追溯),查阅学术论文、技术报告和模型仓库文档,共找到46个模型
- 02在Scopus、IEEEXplore、Web of Science和arXiv四个平台检索到519篇论文,经去重和筛选后先确定32个模型,再通过滚雪球检索追加14个,最终共46个
- 03每个模型按基础模型(BERT、T5、GPT、Llama、Mistral、Gemma、Phi、Qwen等)、用途(通用或特定领域,包括法律、医疗、金融、推特、政府、石油天然气等)、参数量、许可协议和训练数据进行分类
- 04按照11项质量标准对文档完整度打分,平均得分为7.5分(满分11分),其中伦理考量说明(0.07分)、结果的定性分析(0.14分)和研究局限性说明(0.28分)得分最低
- 05按国家统计,巴西以35个模型领先,葡萄牙以10个模型次之;机构方面,巴西公立大学USP(7个)和UNICAMP(6个)领先,企业中Maritaca AI(5个)领先
他们做了什么
- 采用系统性文献梳理方法,结合滚雪球式检索(顺着已知论文的引用关系向前向后追溯),查阅学术论文、技术报告和模型仓库文档,共找到46个模型
- 在Scopus、IEEEXplore、Web of Science和arXiv四个平台检索到519篇论文,经去重和筛选后先确定32个模型,再通过滚雪球检索追加14个,最终共46个
- 每个模型按基础模型(BERT、T5、GPT、Llama、Mistral、Gemma、Phi、Qwen等)、用途(通用或特定领域,包括法律、医疗、金融、推特、政府、石油天然气等)、参数量、许可协议和训练数据进行分类
- 按照11项质量标准对文档完整度打分,平均得分为7.5分(满分11分),其中伦理考量说明(0.07分)、结果的定性分析(0.14分)和研究局限性说明(0.28分)得分最低
- 按国家统计,巴西以35个模型领先,葡萄牙以10个模型次之;机构方面,巴西公立大学USP(7个)和UNICAMP(6个)领先,企业中Maritaca AI(5个)领先
为什么重要
葡萄牙语相较英语等语言资源相对匮乏,其AI模型信息分散在论文、技术报告和代码仓库中,难以形成全局认识,这项研究首次将其系统梳理成图。这为从事葡萄牙语AI研发或应用的开发者和研究者提供了一份清晰的现状地图和空白点参考。
本文术语
- 滚雪球检索(snowballing) · 顺着已知相关论文的引用和被引用关系继续查找更多相关文献的检索方法
- 系统性梳理研究(systematic mapping study) · 按照固定、可重复的流程对某一主题的文献进行调研和分类的研究方法
- 系统发育树(phylogeny) · 以树状图展示哪些模型是基于哪些早期模型构建而来的衍生关系
- 模型卡(Model Card) · 总结模型架构、训练数据和使用方法的说明文档
论文原文摘要(英文)
In recent years, the rapid development of language models has transformed the field of Natural Language Processing through a wide range of applications. However, the development of language models has not progressed uniformly across all languages. In the case of the Portuguese language, there has recently been a growing effort by academia and companies to develop language models and create data resources for Portuguese. These efforts have resulted in the rise of an increasingly diverse ecosystem of language models for Portuguese. However, information on these models remains dispersed in scientific publications, technical reports, model repositories, and project documentation. This survey presents a systematic mapping study of language models developed for Portuguese, providing a comprehensive overview of the current state of the field. We map a total of 46 models, characterizing them by various aspects, including base model, architecture, computational resources, training datasets, licensing, code availability, data, and model weights. Furthermore, we analyzed the evolution and relationships among these models through a phylogenetic perspective, identified current research gaps and opportunities, and discussed future directions for the development of language models for Portuguese.
在 arXiv 阅读最新论文
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?让AI编程助手去修复真实科学软件,连最强的那个也有一半以上任务没做对
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving把稀疏注意力从论文原型变成能真正上线服务的加速方案
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents让客服AI坐席不只是拦住一个危险动作,而是把整个流程走对
- EXIMO: VLM Guided Exploration of VLA Policies不用人工遥控演示,让会说话的AI来教机械臂做新家务
- EnvHarness: Awakening Static Worlds for Agent Learning不重新搭建训练环境,而是给现有环境套一层可插拔组件,针对每个智能体的具体弱点重新塑形
- Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the Model租用AI而非拥有AI的机构,安全监管能力只剩一半
- Beyond Imitation: Filtering On-Policy Distillation by Reasoning ProgressAI模仿老师模型学习时,会误伤本来推理正确的步骤,新方法专门过滤掉这种误伤
- PersonalBench: Measuring the Authorship Gap in LLM Personalization让AI模仿某人的文风,结果发现它始终摆脱不了自己的腔调