FrenchNews-7: Benchmarking Cross-Publisher French News Editorial Desk Classification
一个把法国新闻文章自动分到7个编辑版块的基准测试,换家媒体也照样管用
研究者从13家法国媒体收集了87,637篇文章,直接分析文章网址结构,归纳出社会、文化休闲、国际、政治、体育、经济、科技这7个编辑版块分类体系。用这些数据微调的CamemBERT模型在完全没见过的4家媒体的文章上达到宏观F1值0.799,超过了零样本运行的GPT-OSS-120B、Mistral Small 3.2和Llama-3.3-70B。经济和社会这两类最难分,连人类标注者意见都不统一,说明这更多是媒体编辑习惯不同造成的,而非模型能力问题。
METAL MEDIA 解读图
一个把法国新闻文章自动分到7个编辑版块的基准测试,换家媒体也照样管用
- 01研究者分析了13家法国媒体(如Le Monde、Le Figaro)共87,637篇文章的网址路径片段,据此归纳出扎根于新闻编辑室实际分类习惯的7类版块体系。
- 02对于网址片段含义不明确的27.8%文章,用大语言模型GPT-OSS-120B打标签,再由2名人类和2名大语言模型对整个流程进行核查。
- 03基于法语专门预训练的CamemBERT模型经微调后,在完全未参与训练的4家媒体的2,100篇文章上达到宏观F1值0.799,超过零样本方式的GPT-OSS-120B(0.758)、Mistral Small 3.2(0.775)和Llama-3.3-70B(0.771)。
- 04经济类别表现最弱,召回率只有0.517,这和蒙眼人类标注者在同一批文章上仅55%的一致率相近,说明难点在于不同媒体编辑归类习惯本身就不统一,而不是模型的局限。
- 05研究团队已将微调好的模型和标注数据集发布到Hugging Face,方便其他研究者用于跨媒体多样性分析和议程设置等研究。
他们做了什么
- 研究者分析了13家法国媒体(如Le Monde、Le Figaro)共87,637篇文章的网址路径片段,据此归纳出扎根于新闻编辑室实际分类习惯的7类版块体系。
- 对于网址片段含义不明确的27.8%文章,用大语言模型GPT-OSS-120B打标签,再由2名人类和2名大语言模型对整个流程进行核查。
- 基于法语专门预训练的CamemBERT模型经微调后,在完全未参与训练的4家媒体的2,100篇文章上达到宏观F1值0.799,超过零样本方式的GPT-OSS-120B(0.758)、Mistral Small 3.2(0.775)和Llama-3.3-70B(0.771)。
- 经济类别表现最弱,召回率只有0.517,这和蒙眼人类标注者在同一批文章上仅55%的一致率相近,说明难点在于不同媒体编辑归类习惯本身就不统一,而不是模型的局限。
- 研究团队已将微调好的模型和标注数据集发布到Hugging Face,方便其他研究者用于跨媒体多样性分析和议程设置等研究。

| Family | Articles | % corpus | Publishers |
|---|---|---|---|
| International | 16,090 | 18.0% | 13/13 |
| Culture/Loisirs | 15,637 | 17.5% | 13/13 |
| Politique | 11,513 | 12.9% | 13/13 |
| Société | 11,234 | 12.6% | 13/13 |
| Économie | 8,233 | 9.2% | 13/13 |
| Sport | 7,178 | 8.0% | 13/13 |
| Sciences & Tech. | 3,144 | 3.5% | 10/13 |
| Environnement | 1,976 | 2.2% | 10/13 |
| Santé | 1,030 | 1.2% | 9/13 |
| Médias | 903 | 1.0% | 11/13 |
| Opinion/Édito | 873 | 1.0% | 6/13 |
| Régional | 1,366 | 1.5% | 5/13 |
| Format/Misc | 5,570 | 6.2% | 11/13 |

| Publisher | Articles | % | Type | Bucket A |
|---|---|---|---|---|
| Le Monde | 33,090 | 37.8 | NP | 63.2 |
| L’Express | 9,093 | 10.4 | NP | 82.6 |
| L’Humanité | 7,975 | 9.1 | NP | 68.6 |
| Le Point | 7,877 | 9.0 | NP | 87.6 |
| Le Figaro | 7,270 | 8.3 | NP | 63.0 |
| JDD | 7,269 | 8.3 | NP | 100.0 |
| La Croix | 1,910 | 2.2 | NP | 52.7 |
| Le HuffPost | 2,473 | 2.8 | DN | 84.0 |
| TF1 INFO | 2,044 | 2.3 | BR | 97.4 |
| Le Parisien | 2,375 | 2.7 | NP | 82.9 |
| Slate.fr | 2,105 | 2.4 | DN | 34.3 |
| 20 Minutes | 2,291 | 2.6 | DN | 94.9 |
| Ouest-France | 1,865 | 2.1 | RP | 39.0 |
| Total | 87,637∗ | 100 | 13 |

| Category | Bucket A | Bucket B | Total | % |
|---|---|---|---|---|
| Société | 12,666 | 6,658 | 19,324 | 22.1% |
| International | 15,510 | 3,132 | 18,642 | 21.3% |
| Culture & Loisirs | 13,305 | 5,045 | 18,350 | 20.9% |
| Politique | 7,406 | 3,140 | 10,546 | 12.0% |
| Sport | 6,478 | 1,926 | 8,404 | 9.6% |
| Économie | 5,955 | 2,219 | 8,174 | 9.3% |
| Sciences & Technologies | 1,982 | 2,215 | 4,197 | 4.8% |
| Total | 63,302 | 24,335 | 87,637 | 100% |
| SHA-256 corpus fingerprint: c14a8ae609d009c9289393e6f0be761dda9f015aa4c17f515b09388742a2c218 |

| Model | Context | In-dist. Acc | In-dist. F1 | Cross-pub F1 |
|---|---|---|---|---|
| CamemBERT-base | 512 tok | 0.860 | 0.847 | 0.799 |
| CamemBERTav2 | 512 tok | 0.856 | 0.843 | — |
| CamemBERTav2 | 1,024 tok | 0.861 | 0.847 | 0.798 |
| Model A | Model B | Δ F1 | Bootstrap | McNemar | ||||
|---|---|---|---|---|---|---|---|---|
| p (Holm) | 95% CI | Sig. | χ2 | p (Holm) | Sig. | |||
| CamemBERT-v1 | CamemBERTav2 | +0.004 | 0.086 | [−0.002, 0.009] | 2.11 | 0.147 | ||
| CamemBERT-v1 | mBERT | +0.014 | <0.001 | [0.008, 0.020] | ∗ | 31.26 | <0.001 | ∗ |
| CamemBERT-v1 | TF-IDF | +0.035 | <0.001 | [0.029, 0.042] | ∗ | 152.34 | <0.001 | ∗ |
| CamemBERTav2 | mBERT | +0.010 | 0.001 | [0.005, 0.016] | ∗ | 18.91 | <0.001 | ∗ |
| CamemBERTav2 | TF-IDF | +0.032 | <0.001 | [0.025, 0.038] | ∗ | 128.90 | <0.001 | ∗ |
| mBERT | TF-IDF | +0.021 | <0.001 | [0.015, 0.028] | ∗ | 56.70 | <0.001 | ∗ |
为什么重要
此前没有专门针对跨媒体场景的法语新闻分类基准,这项工作填补了这一空白,为比较不同法国媒体报道倾向的研究提供了实用工具。它也证明了一个体量更小、经过针对性微调的模型,在具体实际任务上可以胜过体量大得多的通用大语言模型。
本文术语
- CamemBERT · 基于约138GB法语文本预训练的RoBERTa结构语言模型
- 网址片段(slug) · 新闻文章网址路径中表示版块名称的部分,例如lemonde.fr/politique/
- 宏观F1(macro-F1) · 对每个类别的表现取相同权重平均后得到的整体准确度指标
- 零样本(zero-shot) · 不给模型任何示例,只凭说明直接让它分类
- 留出集(held-out) · 完全不参与训练、只用来测试模型泛化能力的数据
无法转载的图表
- Figure 5: Top-10 tokens per class by aggregated Integrated Gradients attribution score (199-article stratified sample, CamemBERT-base full-text model). All leading tokens are topically interpretable; no publisher-identifying surface marker appears in any top-10 list.
论文原文摘要(英文)
We present FrenchNews-7, a cross-publisher France-based French-language news editorial desk classification benchmark combining a large multi-outlet corpus, a URL-derived seven-class taxonomy, and a fine-tuned CamemBERT classifier. Labels are assigned via a hybrid pipeline combining publisher URL slugs with LLM annotation for structurally ambiguous cases, audited through an inter-rater study (2 humans + 2 LLMs; pairwise $\kappa \geq 0.766$, human--human $\kappa = 0.806$). We evaluate lexical, multilingual, and French-specific trained classifiers under both in-distribution and held-out-publisher settings, with additional comparison against zero-shot LLM baselines (GPT-OSS-120B, Mistral Small 3.2, Llama-3.3-70B) on the held-out pool. The strongest model, CamemBERT-base on full article text, outperforms headline-only input, generalizes to unseen outlets, and exceeds all three zero-shot LLM baselines on overall recall (0.799), with the gap concentrated in the ambiguous editorial-boundary categories Economie and Societe. Cross-publisher evaluation reveals uneven boundary stability: Sport, Culture & Loisirs, and International transfer cleanly, while Economie (recall = 0.517) is close to blinded human agreement (0.55), and Societe (precision = 0.577) absorbs boundary ambiguity, both suggesting editorial conventions rather than recoverable classifier headroom. The fine-tuned CamemBERT-base model, labeled manifest, reference collection scripts, and a reliability-tier guidance table are available at https://huggingface.co/LeFrenchNewsLab/camembert-base-frenchnews7 (model) and https://huggingface.co/datasets/LeFrenchNewsLab/frenchnews-7 (dataset).
在 arXiv 阅读最新论文
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?让AI编程助手去修复真实科学软件,连最强的那个也有一半以上任务没做对
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving把稀疏注意力从论文原型变成能真正上线服务的加速方案
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents让客服AI坐席不只是拦住一个危险动作,而是把整个流程走对
- EXIMO: VLM Guided Exploration of VLA Policies不用人工遥控演示,让会说话的AI来教机械臂做新家务
- EnvHarness: Awakening Static Worlds for Agent Learning不重新搭建训练环境,而是给现有环境套一层可插拔组件,针对每个智能体的具体弱点重新塑形
- Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the Model租用AI而非拥有AI的机构,安全监管能力只剩一半
- Beyond Imitation: Filtering On-Policy Distillation by Reasoning ProgressAI模仿老师模型学习时,会误伤本来推理正确的步骤,新方法专门过滤掉这种误伤
- PersonalBench: Measuring the Authorship Gap in LLM Personalization让AI模仿某人的文风,结果发现它始终摆脱不了自己的腔调
METAL MEDIA 最新报道
图片来源: Amr Sobhy et al., arXiv:2608.18097, CC BY 4.0