컴백부터 K-뷰티까지 — K-컬쳐의 모든 것을 메일로 받아보세요메일로 받아보기

METAL MEDIA

프랑스 뉴스 기사를 7개 편집국 카테고리로 자동 분류하는 벤치마크, 신문사가 바뀌어도 꽤 잘 통한다

arXiv:2608.180972026-08-20

FrenchNews-7: Benchmarking Cross-Publisher French News Editorial Desk Classification

프랑스 뉴스 기사를 7개 편집국 카테고리로 자동 분류하는 벤치마크, 신문사가 바뀌어도 꽤 잘 통한다

연구자들은 13개 프랑스 언론사에서 87,637개 기사를 모아 URL 구조를 분석해 사회, 문화·여가, 국제, 정치, 스포츠, 경제, 과학·기술 등 7개 편집국 카테고리 체계를 만들었다. 프랑스어 특화 언어모델 CamemBERT를 이 데이터로 학습시켰더니 학습에 전혀 쓰지 않은 4개 언론사 기사에서도 macro-F1 0.799를 기록하며 대형 LLM(GPT-OSS-120B, Mistral, Llama-3.3-70B)보다 높은 재현율을 보였다. 다만 경제와 사회 카테고리는 신문사마다 편집 관행이 달라 사람도 헷갈리는 경계 영역으로 확인됐다.

METAL MEDIA 해설 도표

프랑스 뉴스 기사를 7개 편집국 카테고리로 자동 분류하는 벤치마크, 신문사가 바뀌어도 꽤 잘 통한다

  1. 01프랑스 13개 언론사(르몽드, 르피가로, 위마니테 등) 기사 87,637개를 모아 기사 URL의 섹션 이름(슬러그)을 분석해 7개 편집국 카테고리 체계를 설계했다.
  2. 02URL만으로 분류가 애매한 기사(전체의 27.8%)는 대형언어모델(GPT-OSS-120B)에게 맡겨 라벨을 붙이고, 사람 2명과 LLM 2명이 검증해 신뢰도를 확인했다.
  3. 03프랑스어 전용으로 학습된 CamemBERT 모델을 미세조정(fine-tuning)한 결과, 처음 보는 4개 언론사 기사 2,100개에서 macro-F1 0.799를 달성해 제로샷(사례 없이 바로 분류)으로 돌린 GPT-OSS-120B(0.758), Mistral Small 3.2(0.775), Llama-3.3-70B(0.771)를 모두 앞질렀다.
  4. 04경제 카테고리는 재현율 0.517로 가장 약했는데, 이는 사람이 봐도 신문사별 경제 기사 분류 기준이 제각각이라 55% 정도밖에 일치하지 않는 것과 비슷한 수준이라 모델의 한계라기보다 언론사 편집 관행 차이로 해석된다.
  5. 05학습된 모델과 라벨 데이터셋을 Hugging Face에 공개해 다른 연구자들이 프랑스 언론 다양성 분석, 의제 설정 연구 등에 바로 활용할 수 있게 했다.
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.

무엇을 했나

  1. 프랑스 13개 언론사(르몽드, 르피가로, 위마니테 등) 기사 87,637개를 모아 기사 URL의 섹션 이름(슬러그)을 분석해 7개 편집국 카테고리 체계를 설계했다.
  2. URL만으로 분류가 애매한 기사(전체의 27.8%)는 대형언어모델(GPT-OSS-120B)에게 맡겨 라벨을 붙이고, 사람 2명과 LLM 2명이 검증해 신뢰도를 확인했다.
  3. 프랑스어 전용으로 학습된 CamemBERT 모델을 미세조정(fine-tuning)한 결과, 처음 보는 4개 언론사 기사 2,100개에서 macro-F1 0.799를 달성해 제로샷(사례 없이 바로 분류)으로 돌린 GPT-OSS-120B(0.758), Mistral Small 3.2(0.775), Llama-3.3-70B(0.771)를 모두 앞질렀다.
  4. 경제 카테고리는 재현율 0.517로 가장 약했는데, 이는 사람이 봐도 신문사별 경제 기사 분류 기준이 제각각이라 55% 정도밖에 일치하지 않는 것과 비슷한 수준이라 모델의 한계라기보다 언론사 편집 관행 차이로 해석된다.
  5. 학습된 모델과 라벨 데이터셋을 Hugging Face에 공개해 다른 연구자들이 프랑스 언론 다양성 분석, 의제 설정 연구 등에 바로 활용할 수 있게 했다.
Figure 1: Row-normalized confusion matrix for CamemBERT-base on the in-distribution test set (n=13,146).
Figure 1: Row-normalized confusion matrix for CamemBERT-base on the in-distribution test set (n=13,146).
Table 1: Candidate URL families ranked by volume. The rule marks the seven-family break; lower families are absorbed into semantic parents or deferred to Bucket B. The 2,890 articles with absent or unparseable URL paths are labeled via Bucket B.
FamilyArticles% corpusPublishers
International16,09018.0%13/13
Culture/Loisirs15,63717.5%13/13
Politique11,51312.9%13/13
Société11,23412.6%13/13
Économie8,2339.2%13/13
Sport7,1788.0%13/13
Sciences & Tech.3,1443.5%10/13
Environnement1,9762.2%10/13
Santé1,0301.2%9/13
Médias9031.0%11/13
Opinion/Édito8731.0%6/13
Régional1,3661.5%5/13
Format/Misc5,5706.2%11/13
Figure 2: Per-class precision, recall, and F1 on the pooled mixed-class cross-publisher test set (n=2,100, 300 per category).
Figure 2: Per-class precision, recall, and F1 on the pooled mixed-class cross-publisher test set (n=2,100, 300 per category).
Table 2: Corpus composition by publisher. Type: NP = national press, RP = regional press, DN = digital-native, BR = broadcast web. Bucket A = deterministic slug-label share. ∗Post-deduplication; 132 duplicates removed.
PublisherArticles%TypeBucket A
Le Monde33,09037.8NP63.2
L’Express9,09310.4NP82.6
L’Humanité7,9759.1NP68.6
Le Point7,8779.0NP87.6
Le Figaro7,2708.3NP63.0
JDD7,2698.3NP100.0
La Croix1,9102.2NP52.7
Le HuffPost2,4732.8DN84.0
TF1 INFO2,0442.3BR97.4
Le Parisien2,3752.7NP82.9
Slate.fr2,1052.4DN34.3
20 Minutes2,2912.6DN94.9
Ouest-France1,8652.1RP39.0
Total87,637∗10013
Figure 3: Editorial topic fingerprints (2022–2025). Predicted topic shares for 13 outlets across 7 categories, aggregated from the unlabelled longitudinal corpus. Each cell shows the percentage of articles assigned to a given topic; rows sum to 100%.
Figure 3: Editorial topic fingerprints (2022–2025). Predicted topic shares for 13 outlets across 7 categories, aggregated from the unlabelled longitudinal corpus. Each cell shows the percentage of articles assigned to a given topic; rows sum to 100%.
Table 3: Dataset distribution by label source. Bucket A = deterministic slug rules; Bucket B = LLM-labeled ambiguous slugs.
CategoryBucket ABucket BTotal%
Société12,6666,65819,32422.1%
International15,5103,13218,64221.3%
Culture & Loisirs13,3055,04518,35020.9%
Politique7,4063,14010,54612.0%
Sport6,4781,9268,4049.6%
Économie5,9552,2198,1749.3%
Sciences & Technologies1,9822,2154,1974.8%
Total63,30224,33587,637100%
SHA-256 corpus fingerprint: c14a8ae609d009c9289393e6f0be761dda9f015aa4c17f515b09388742a2c218
Figure 4: Pairwise Jensen–Shannon divergence of outlet topic distributions over 48 months (2022–2025). Higher values indicate greater topical differentiation (see text for statistics). All values are computed from classifier predictions, not gold labels.
Figure 4: Pairwise Jensen–Shannon divergence of outlet topic distributions over 48 months (2022–2025). Higher values indicate greater topical differentiation (see text for statistics). All values are computed from classifier predictions, not gold labels.
Table 12: Comparison of CamemBERT-base (512 tokens) with CamemBERTav2 at 512-token truncation and its native 1,024-token context. In-distribution metrics on the 13,146-article test split; cross-publisher macro-F1 on the 2,100-article held-out outlet pool.
ModelContextIn-dist. AccIn-dist. F1Cross-pub F1
CamemBERT-base512 tok0.8600.8470.799
CamemBERTav2512 tok0.8560.843
CamemBERTav21,024 tok0.8610.8470.798
Table 13: In-distribution pairwise significance (13,146 test set, seed 42). Paired bootstrap (10,000 replicates) on macro-F1; McNemar’s test on per-article correctness. Holm correction over all 12 raw p-values. Δ = macro-F1(A) − macro-F1(B). p∗<0.05 after correction.
Model AModel BΔ F1BootstrapMcNemar
p (Holm)95% CISig.χ2p (Holm)Sig.
CamemBERT-v1CamemBERTav2+0.0040.086[−0.002, 0.009]2.110.147
CamemBERT-v1mBERT+0.014<0.001[0.008, 0.020]31.26<0.001
CamemBERT-v1TF-IDF+0.035<0.001[0.029, 0.042]152.34<0.001
CamemBERTav2mBERT+0.0100.001[0.005, 0.016]18.91<0.001
CamemBERTav2TF-IDF+0.032<0.001[0.025, 0.038]128.90<0.001
mBERTTF-IDF+0.021<0.001[0.015, 0.028]56.70<0.001

왜 중요한가

프랑스어 뉴스에서 언론사를 넘나드는 자동 분류기가 없었던 공백을 채워, 여러 매체의 보도 경향을 한꺼번에 비교하는 미디어 연구에 실질적인 도구를 제공한다. 대형 범용 LLM보다 작고 특화된 모델이 특정 실무 과제에서 더 낫다는 것을 실증한 사례이기도 하다.

이 논문의 용어

  • CamemBERT · 프랑스어 텍스트 약 138GB로 사전학습된 프랑스어 전용 언어모델
  • 슬러그(slug) · 기사 URL 경로에 포함된 섹션 이름 조각, 예: lemonde.fr/politique/
  • macro-F1 · 각 카테고리별 정확도를 동일 가중치로 평균낸 전체 성능 지표
  • 제로샷(zero-shot) · 별도 예시 없이 설명만 주고 바로 분류를 시키는 방식
  • held-out(홀드아웃) · 모델 학습에 전혀 쓰이지 않고 평가에만 사용하는 데이터

본문에 싣지 못한 그림

  • Figure 5: Top-10 tokens per class by aggregated Integrated Gradients attribution score (199-article stratified sample, CamemBERT-base full-text model). All leading tokens are topically interpretable; no publisher-identifying surface marker appears in any top-10 list.
원문에서 그림 보기 →

저자 · Amr Sobhy

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL MEDIA 최신 기사

그림 출처: Amr Sobhy et al., arXiv:2608.18097, CC BY 4.0