Everything K-culture — comebacks to K-beauty, straight to your inboxGet it in your inbox

METAL MEDIA

FrenchNews-7: Benchmarking Cross-Publisher French News Editorial Desk Classification

arXiv:2608.180972026-08-20

A benchmark that sorts French news articles into 7 editorial desks and holds up even on outlets it never saw

Researchers built FrenchNews-7 from 87,637 articles across 13 France-based outlets, deriving a seven-category editorial desk taxonomy (Société, Culture & Loisirs, International, Politique, Sport, Économie, Sciences & Technologies) directly from article URL structure. A fine-tuned CamemBERT model reached macro-F1 0.799 on four completely unseen outlets, beating zero-shot GPT-OSS-120B, Mistral Small 3.2, and Llama-3.3-70B. The economy and society categories remained the hardest, matching how much even humans disagree on where those articles belong editorially.

METAL MEDIA explanatory visual

A benchmark that sorts French news articles into 7 editorial desks and holds up even on outlets it never saw

  1. 01The authors analyzed URL path segments (slugs) from 87,637 articles across 13 French outlets like Le Monde and Le Figaro to build a 7-class editorial taxonomy grounded in how newsrooms actually organize their sections.
  2. 02For the 27.8% of articles whose URL slug was ambiguous, an LLM (GPT-OSS-120B) assigned labels, and the whole pipeline was audited with 2 human raters and 2 LLM raters for reliability.
  3. 03A CamemBERT model (a RoBERTa-style encoder pretrained specifically on French text) was fine-tuned on this data and achieved macro-F1 0.799 on 2,100 articles from four outlets never used in training, outperforming zero-shot LLM baselines GPT-OSS-120B (0.758), Mistral Small 3.2 (0.775), and Llama-3.3-70B (0.771).
  4. 04The Économie category was weakest, with recall of only 0.517 — close to the 55% agreement rate blinded human annotators showed on the same articles, suggesting the difficulty reflects genuinely inconsistent editorial conventions across publishers rather than a model shortcoming.
  5. 05The fine-tuned model and labeled dataset were released on Hugging Face so other researchers can use them for cross-outlet media diversity studies and agenda-setting research.
An explanatory diagram made by METAL MEDIA, not a figure supplied by the paper's authors.

What they did

  1. The authors analyzed URL path segments (slugs) from 87,637 articles across 13 French outlets like Le Monde and Le Figaro to build a 7-class editorial taxonomy grounded in how newsrooms actually organize their sections.
  2. For the 27.8% of articles whose URL slug was ambiguous, an LLM (GPT-OSS-120B) assigned labels, and the whole pipeline was audited with 2 human raters and 2 LLM raters for reliability.
  3. A CamemBERT model (a RoBERTa-style encoder pretrained specifically on French text) was fine-tuned on this data and achieved macro-F1 0.799 on 2,100 articles from four outlets never used in training, outperforming zero-shot LLM baselines GPT-OSS-120B (0.758), Mistral Small 3.2 (0.775), and Llama-3.3-70B (0.771).
  4. The Économie category was weakest, with recall of only 0.517 — close to the 55% agreement rate blinded human annotators showed on the same articles, suggesting the difficulty reflects genuinely inconsistent editorial conventions across publishers rather than a model shortcoming.
  5. The fine-tuned model and labeled dataset were released on Hugging Face so other researchers can use them for cross-outlet media diversity studies and agenda-setting research.
Figure 1: Row-normalized confusion matrix for CamemBERT-base on the in-distribution test set (n=13,146).
Figure 1: Row-normalized confusion matrix for CamemBERT-base on the in-distribution test set (n=13,146).
Table 1: Candidate URL families ranked by volume. The rule marks the seven-family break; lower families are absorbed into semantic parents or deferred to Bucket B. The 2,890 articles with absent or unparseable URL paths are labeled via Bucket B.
FamilyArticles% corpusPublishers
International16,09018.0%13/13
Culture/Loisirs15,63717.5%13/13
Politique11,51312.9%13/13
Société11,23412.6%13/13
Économie8,2339.2%13/13
Sport7,1788.0%13/13
Sciences & Tech.3,1443.5%10/13
Environnement1,9762.2%10/13
Santé1,0301.2%9/13
Médias9031.0%11/13
Opinion/Édito8731.0%6/13
Régional1,3661.5%5/13
Format/Misc5,5706.2%11/13
Figure 2: Per-class precision, recall, and F1 on the pooled mixed-class cross-publisher test set (n=2,100, 300 per category).
Figure 2: Per-class precision, recall, and F1 on the pooled mixed-class cross-publisher test set (n=2,100, 300 per category).
Table 2: Corpus composition by publisher. Type: NP = national press, RP = regional press, DN = digital-native, BR = broadcast web. Bucket A = deterministic slug-label share. ∗Post-deduplication; 132 duplicates removed.
PublisherArticles%TypeBucket A
Le Monde33,09037.8NP63.2
L’Express9,09310.4NP82.6
L’Humanité7,9759.1NP68.6
Le Point7,8779.0NP87.6
Le Figaro7,2708.3NP63.0
JDD7,2698.3NP100.0
La Croix1,9102.2NP52.7
Le HuffPost2,4732.8DN84.0
TF1 INFO2,0442.3BR97.4
Le Parisien2,3752.7NP82.9
Slate.fr2,1052.4DN34.3
20 Minutes2,2912.6DN94.9
Ouest-France1,8652.1RP39.0
Total87,637∗10013
Figure 3: Editorial topic fingerprints (2022–2025). Predicted topic shares for 13 outlets across 7 categories, aggregated from the unlabelled longitudinal corpus. Each cell shows the percentage of articles assigned to a given topic; rows sum to 100%.
Figure 3: Editorial topic fingerprints (2022–2025). Predicted topic shares for 13 outlets across 7 categories, aggregated from the unlabelled longitudinal corpus. Each cell shows the percentage of articles assigned to a given topic; rows sum to 100%.
Table 3: Dataset distribution by label source. Bucket A = deterministic slug rules; Bucket B = LLM-labeled ambiguous slugs.
CategoryBucket ABucket BTotal%
Société12,6666,65819,32422.1%
International15,5103,13218,64221.3%
Culture & Loisirs13,3055,04518,35020.9%
Politique7,4063,14010,54612.0%
Sport6,4781,9268,4049.6%
Économie5,9552,2198,1749.3%
Sciences & Technologies1,9822,2154,1974.8%
Total63,30224,33587,637100%
SHA-256 corpus fingerprint: c14a8ae609d009c9289393e6f0be761dda9f015aa4c17f515b09388742a2c218
Figure 4: Pairwise Jensen–Shannon divergence of outlet topic distributions over 48 months (2022–2025). Higher values indicate greater topical differentiation (see text for statistics). All values are computed from classifier predictions, not gold labels.
Figure 4: Pairwise Jensen–Shannon divergence of outlet topic distributions over 48 months (2022–2025). Higher values indicate greater topical differentiation (see text for statistics). All values are computed from classifier predictions, not gold labels.
Table 12: Comparison of CamemBERT-base (512 tokens) with CamemBERTav2 at 512-token truncation and its native 1,024-token context. In-distribution metrics on the 13,146-article test split; cross-publisher macro-F1 on the 2,100-article held-out outlet pool.
ModelContextIn-dist. AccIn-dist. F1Cross-pub F1
CamemBERT-base512 tok0.8600.8470.799
CamemBERTav2512 tok0.8560.843
CamemBERTav21,024 tok0.8610.8470.798
Table 13: In-distribution pairwise significance (13,146 test set, seed 42). Paired bootstrap (10,000 replicates) on macro-F1; McNemar’s test on per-article correctness. Holm correction over all 12 raw p-values. Δ = macro-F1(A) − macro-F1(B). p∗<0.05 after correction.
Model AModel BΔ F1BootstrapMcNemar
p (Holm)95% CISig.χ2p (Holm)Sig.
CamemBERT-v1CamemBERTav2+0.0040.086[−0.002, 0.009]2.110.147
CamemBERT-v1mBERT+0.014<0.001[0.008, 0.020]31.26<0.001
CamemBERT-v1TF-IDF+0.035<0.001[0.029, 0.042]152.34<0.001
CamemBERTav2mBERT+0.0100.001[0.005, 0.016]18.91<0.001
CamemBERTav2TF-IDF+0.032<0.001[0.025, 0.038]128.90<0.001
mBERTTF-IDF+0.021<0.001[0.015, 0.028]56.70<0.001

Why it matters

This fills a gap for French-language media research by providing a classifier that reliably generalizes across publishers, something no prior French benchmark directly targeted. It also shows a smaller, specialized fine-tuned model can beat much larger general-purpose LLMs on a focused real-world classification task.

Terms in this paper

  • CamemBERT · A RoBERTa-based language model pretrained on roughly 138GB of French text
  • slug · The section-name segment embedded in a news article's URL path, e.g. lemonde.fr/politique/
  • macro-F1 · An overall accuracy score that averages performance equally across all categories
  • zero-shot · Asking a model to classify without giving it any labeled examples first
  • held-out · Data kept completely separate from training, used only to test generalization

Figures we cannot republish

  • Figure 5: Top-10 tokens per class by aggregated Integrated Gradients attribution score (199-article stratified sample, CamemBERT-base full-text model). All leading tokens are topically interpretable; no publisher-identifying surface marker appears in any top-10 list.
See the figures in the original paper →

Original abstract (English)

We present FrenchNews-7, a cross-publisher France-based French-language news editorial desk classification benchmark combining a large multi-outlet corpus, a URL-derived seven-class taxonomy, and a fine-tuned CamemBERT classifier. Labels are assigned via a hybrid pipeline combining publisher URL slugs with LLM annotation for structurally ambiguous cases, audited through an inter-rater study (2 humans + 2 LLMs; pairwise $\kappa \geq 0.766$, human--human $\kappa = 0.806$). We evaluate lexical, multilingual, and French-specific trained classifiers under both in-distribution and held-out-publisher settings, with additional comparison against zero-shot LLM baselines (GPT-OSS-120B, Mistral Small 3.2, Llama-3.3-70B) on the held-out pool. The strongest model, CamemBERT-base on full article text, outperforms headline-only input, generalizes to unseen outlets, and exceeds all three zero-shot LLM baselines on overall recall (0.799), with the gap concentrated in the ambiguous editorial-boundary categories Economie and Societe. Cross-publisher evaluation reveals uneven boundary stability: Sport, Culture & Loisirs, and International transfer cleanly, while Economie (recall = 0.517) is close to blinded human agreement (0.55), and Societe (precision = 0.577) absorbs boundary ambiguity, both suggesting editorial conventions rather than recoverable classifier headroom. The fine-tuned CamemBERT-base model, labeled manifest, reference collection scripts, and a reliability-tier guidance table are available at https://huggingface.co/LeFrenchNewsLab/camembert-base-frenchnews7 (model) and https://huggingface.co/datasets/LeFrenchNewsLab/frenchnews-7 (dataset).

Authors · Amr Sobhy

Read on arXiv

Latest papers

All papers →

Latest from METAL MEDIA

Figures: Amr Sobhy et al., arXiv:2608.18097, CC BY 4.0