긴 소설 요약할 때 AI가 지어내는 거짓말, 얼마나 잘 잡아낼까 시험하는 벤치마크가 나왔다
arXiv:2608.180822026-08-20
LongNovel: A Multi-Scale Benchmark for Hallucination Detection in Long-Context Novel Summarization
긴 소설 요약할 때 AI가 지어내는 거짓말, 얼마나 잘 잡아낼까 시험하는 벤치마크가 나왔다
AI 모델이 처리할 수 있는 글자 수(컨텍스트)가 크게 늘었지만, 긴 글을 요약할 때 사실과 다른 내용을 지어내는 '환각' 문제는 여전하다. 연구팀은 중국 소설 29권과 영어 소설 데이터를 이용해 16k부터 100k 토큰까지 네 가지 길이, 8가지 환각 유형을 아우르는 LongNovel 벤치마크를 만들었다. GPT, Claude, DeepSeek 등 여러 최신 모델을 시험한 결과 글이 길어질수록 성능이 눈에 띄게 떨어졌다.
METAL MEDIA 해설 도표
긴 소설 요약할 때 AI가 지어내는 거짓말, 얼마나 잘 잡아낼까 시험하는 벤치마크가 나왔다
01중국어 소설 29권과 영어 소설 모음(BookSum)을 이용해 16k, 32k, 64k, 100k 토큰(글자 분량 단위) 짜리 4단계 길이의 요약-검증 데이터 6,354개를 만들었다.
02실제 AI가 요약하면서 저지르는 자연스러운 실수를 잡아내는 방식(여러 모델이 서로 채점하는 Multi-Model Arbitration)과, 사람이 쓴 요약에 일부러 오류를 심는 방식(Entity-Referenced Hallucination Generation)을 함께 써서 데이터의 사실성과 유형별 균형을 모두 확보했다.
03인물·숫자·관계·사건·시간·인과관계 등 8가지 환각 유형으로 나누고, 테스트셋 204개 요약은 사람이 직접 검토해 신뢰도를 높였다(검토자 간 일치도 0.918).
04GPT-5.2-chat, Claude-4.5-Sonnet, DeepSeek 시리즈, Qwen3 등 오픈소스·상용 모델을 시험한 결과, 글이 길어질수록(16k→100k) 정확도가 일관되게 떨어졌고(예: Claude-4.5-Sonnet 0.755→0.670), 시간 순서·인과관계 오류는 특히 잡아내기 어려워했다.
05긴 글을 조각내 검색하며 검토하는 방식(RAG)과 여러 모델의 판단을 투표로 합치는 방식, 그리고 모델을 직접 재학습(SFT)시키는 방식을 비교했더니, 재학습이 긴 글에서 가장 안정적으로 성능을 끌어올렸다(Qwen3-32B가 100k에서 0.580→0.720).
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.
무엇을 했나
중국어 소설 29권과 영어 소설 모음(BookSum)을 이용해 16k, 32k, 64k, 100k 토큰(글자 분량 단위) 짜리 4단계 길이의 요약-검증 데이터 6,354개를 만들었다.
실제 AI가 요약하면서 저지르는 자연스러운 실수를 잡아내는 방식(여러 모델이 서로 채점하는 Multi-Model Arbitration)과, 사람이 쓴 요약에 일부러 오류를 심는 방식(Entity-Referenced Hallucination Generation)을 함께 써서 데이터의 사실성과 유형별 균형을 모두 확보했다.
인물·숫자·관계·사건·시간·인과관계 등 8가지 환각 유형으로 나누고, 테스트셋 204개 요약은 사람이 직접 검토해 신뢰도를 높였다(검토자 간 일치도 0.918).
GPT-5.2-chat, Claude-4.5-Sonnet, DeepSeek 시리즈, Qwen3 등 오픈소스·상용 모델을 시험한 결과, 글이 길어질수록(16k→100k) 정확도가 일관되게 떨어졌고(예: Claude-4.5-Sonnet 0.755→0.670), 시간 순서·인과관계 오류는 특히 잡아내기 어려워했다.
긴 글을 조각내 검색하며 검토하는 방식(RAG)과 여러 모델의 판단을 투표로 합치는 방식, 그리고 모델을 직접 재학습(SFT)시키는 방식을 비교했더니, 재학습이 긴 글에서 가장 안정적으로 성능을 끌어올렸다(Qwen3-32B가 100k에서 0.580→0.720).
Table 1: Comparison of our benchmark with other benchmarks in novel. ‘Summ. Halluc.’, ‘100K’, ‘Auto. Gen.’, and ‘Diff. Len.’ mean whether it is a summarization hallucination detection dataset, whether it reaches up to 100K tokens, whether the hallucinated data is generated through automated methods, and whether it encompasses different levels of length, respectively. LCHD Liu et al. (2025) refers to the long-context hallucination detection dataset.
Benchmark
Summ. Halluc.
100K
Auto. Label
Diff. Len.
Nocha (2024a)
✗
✓
✗
✗
StorySumm (2024)
✓
✗
✗
✗
FABLES (2024)
✓
✓
✗
✗
LCHD (2025)
✓
✗
✓
✗
CLIPPER (2025)
✗
✓
✓
✓
LongNovel (Ours)
✓
✓
✓
✓
Table 2: Statistics of the LongNovel training and test sets. Hallu. and Non-H. represent the counts of hallucinated and non-hallucinated samples, respectively.
Context
Train
Test
Length
Hallu.
Non-H.
Total
Hallu.
Non-H.
Total
S
1,454
1,430
2,884
100
100
200
M
1,292
1,178
2,470
100
100
200
L
–
–
–
50
50
100
XL
–
–
–
50
50
100
Total
5,354
600
Table 3: Token sequence length of LongNovel test set (values are in thousands, i.e., k).
Tokenizer
S (n=200)
M (n=200)
L (n=100)
XL (n=100)
Min
Mean
Max
Min
Mean
Max
Min
Mean
Max
Min
Mean
Max
GPT-4o
14.65
21.63
33.35
30.15
43.81
60.93
62.34
87.32
119.87
98.78
139.68
187.36
InternLM-2.5
13.37
15.86
18.01
29.31
32.21
34.35
59.62
64.03
66.98
92.81
100.68
104.17
Qwen-3
14.40
15.81
18.06
30.23
32.09
34.48
62.22
63.76
65.93
98.30
100.21
103.32
GLM-4
13.67
15.49
17.62
28.65
31.48
33.58
59.47
62.43
64.91
93.89
98.18
100.70
Table 4: Balanced Accuracy and MCC of various models. The best score is bold and the second-best is underline within each category (Open-Source vs Commercial). Voting Ensembles are abbreviation-coded as follows: (DSV3/DSR/C) denotes DeepSeek-V3, DeepSeek-R1, and Claude-sonnet-4-20250514-v1; (DSV4/GPT/C) denotes DeepSeek-V4, GPT-5.2-chat, and Claude-sonnet-4-20250514-v1.
Model
S(16k)
M(32k)
L(64k)
XL(100k)
BAcc.
MCC.
BAcc.
MCC.
BAcc.
MCC.
BAcc.
MCC.
Open-Source Models
InternLM2.5-20B
0.545
0.217
0.505
0.024
0.450
-0.196
0.480
-0.143
—CoT
0.555
0.241
0.520
0.084
0.500
0.000
0.490
-0.101
Minicheck-7B
0.500
0.000
0.392
-0.335
0.420
-0.266
0.310
-0.484
GLM-4-9B-chat
0.505
0.014
0.510
0.059
0.500
0.000
0.500
0.000
—CoT
0.520
0.048
0.445
-0.139
0.470
-0.094
0.510
0.059
Llama3.1-8B-instruct
0.525
0.052
0.077
-0.837
0.000
-1.000
0.000
-1.000
—CoT
0.418
-0.164
0.042
-0.917
0.000
-1.000
0.000
-1.000
—SFT
0.715
0.508
0.542
0.173
0.210
-0.583
0.310
-0.583
Qwen3-8B
0.568
0.233
0.608
0.237
0.533
0.075
0.480
-0.050
—CoT
0.537
0.131
0.568
0.140
0.577
0.192
0.450
-0.115
—SFT
0.773
0.596
0.770
0.563
0.800
0.643
0.760
0.579
Qwen3-14B
0.588
0.285
0.598
0.219
0.547
0.105
0.517
0.036
—CoT
0.595
0.324
0.573
0.171
0.570
0.167
0.533
0.073
Qwen3-32B
0.590
0.292
0.575
0.180
0.597
0.253
0.580
0.187
—CoT
0.595
0.298
0.577
0.176
0.580
0.215
0.580
0.178
—SFT
0.775
0.587
0.795
0.625
0.750
0.551
0.720
0.490
Commercial Models
Claude-4.5-Sonnet
0.755
0.510
0.665
0.332
0.680
0.363
0.670
0.340
—CoT
0.780
0.564
0.715
0.435
0.670
0.341
0.750
0.503
DeepSeek-v3
0.755
0.544
0.735
0.521
0.680
0.450
0.670
0.433
—CoT
0.735
0.510
0.685
0.424
0.620
0.312
0.610
0.280
—RAG
–
–
–
–
0.530
0.068
0.500
-0.014
DeepSeek-r1
0.775
0.592
0.735
0.510
0.640
0.382
0.680
0.450
—CoT
0.710
0.440
0.680
0.435
0.610
0.327
0.670
0.433
DeepSeek-v4-flash
0.785
0.573
0.810
0.625
0.830
0.660
0.830
0.660
GPT-5.2-chat
0.785
0.570
0.770
0.541
0.810
0.620
0.740
0.482
—CoT
0.775
0.555
0.765
0.530
0.780
0.560
0.750
0.503
—RAG
–
–
–
–
0.830
0.661
0.750
0.500
Table 5: Ablation study on data construction strategies across various models and dataset scales.
Model
S(16k)
M(32k)
L(64k)
XL(100k)
BAcc.
MCC.
BAcc.
MCC.
BAcc.
MCC.
BAcc.
MCC.
Qwen3-8B
0.568
0.233
0.608
0.237
0.533
0.075
0.480
-0.050
Qwen3-8B-MMA
0.577
0.248
0.577
0.171
0.610
0.261
0.710
0.447
Qwen3-8B-FULL
0.773
0.596
0.770
0.563
0.800
0.643
0.760
0.579
Llama3.1-8B-instruct
0.525
0.052
0.082
-0.837
0.000
-1.000
0.000
-1.000
Llama3.1-8B-MMA
0.585
0.290
0.502
0.008
0.020
-0.961
0.000
-1.000
Llama3.1-8B-FULL
0.720
0.508
0.542
0.173
0.210
-0.583
0.310
-0.583
Qwen3-32B
0.590
0.292
0.575
0.180
0.597
0.253
0.580
0.187
Qwen3-32B-MMA
0.605
0.343
0.587
0.281
0.610
0.352
0.680
0.469
Qwen3-32B-FULL
0.775
0.587
0.795
0.625
0.750
0.551
0.720
0.490
Table 6: Statistical distribution and composition of the LongNovel dataset across various subsets. ZH and EN denote Chinese and English data; Original denotes concatenated and compressed human-written summaries; ERHG denotes Entity-Referenced Hallucination Generation; MMA denotes Multi-Model Arbitration; Hallu. and Non-H. represent the counts of hallucinated and non-hallucinated samples
Subset
Total
ZH
EN
Original
ERHG
MMA
Hallu.
Non-H.
Train_16k
2884
1230
1654
814
963
1107
1454
1430
Train_32k
2470
1091
1379
734
937
799
1292
1178
Valid_16k
200
100
100
60
40
100
100
100
Valid_32k
200
100
100
60
40
100
100
100
Test_16k
200
100
100
64
48
88
100
100
Test_32k
200
100
100
69
61
70
100
100
Test_64k
100
50
50
41
27
32
50
50
Test_100k
100
54
46
38
20
42
50
50
Table 7: Detailed composition of the Chinese subset in the LongNovel dataset, listing the book titles and authors across the training, validation, and test splits.
Split
Book Title
Author
Training Set
Bright Eyes in the Dark
Er Dong Tu Zi
Destined
Mo Shu Bai
Misty Rain Tower
Yixi Yanyu
The Gentlemen of the City
Jin Shisichai
We Love Each Other So Much (Chinese edition)
Marcela Serrano
Little Confucian Immortal Seeking the Unknown
Youziyin
Turing’s Code
Feitian Yexiang
Folding Beijing
Hao Jingfang
I’m Waiting for You in the Memory
Xin Yiwu
Bu Yi Jian
Feng Diuzi
The Chronicle of Qingxi: Volume 1
Hong Zhuxia
The Chronicle of Qingxi: Volume 2
Hong Zhuxia
The Chronicle of Qingxi: Volume 3
Hong Zhuxia
The Space-Time Painter
Hai Ya
Once Gone
Qingshan Huangzhong
Validation Set
Yuhong
Ban Yu
Let Me Look at You
Xin Yiwu
I’m Waiting for You in the Memory
Xin Yiwu
My Garden
Teng Ping
Test Set
Bad Kids
Zijin Chen
Exclusive Possession
Ding Mo
Blood is Burning
Bainian Ruge
Everyone is a Protagonist Except Me
Cong Wen
The System Granted Me Longevity
Zi Ling Feng Xue
Mysterious Lotus Casebook
Teng Ping
Cicadas Sing the Setting Sun West
Yu Luo Zhu Leng
Ge Lu Ming: Volume 1
Qin Huai
Ge Lu Ming: Volume 2
Qin Huai
The Creatures That We Are
Peng Pai
Table 8: Detailed performance statistics in 16k, 32k scales. PromptB indicates that the target summary is positioned at the Beginning, while the default setup uses Prompt-E, which positions the target summary at the End. Voting Ensembles are abbreviation-coded as follows: (DSV3/DSR/C) denotes DeepSeek-V3, DeepSeek-R1, and Claude-sonnet-4-20250514-v1; (DSV4/GPT/C) denotes DeepSeek-V4, GPT-5.2-chat, and Claude-sonnet-4-20250514-v1.
Model
S(16k)
M(32k)
Non-H
Hallu
BAcc.
MCC.
Non-H
Hallu
BAcc.
MCC.
Open-Source Models
InternLM2.5-20B
1.000
0.090
0.545
0.217
0.960
0.050
0.505
0.024
—CoT
1.000
0.110
0.555
0.241
0.960
0.080
0.520
0.084
—Prompt-B
0.790
0.380
0.585
0.186
0.950
0.030
0.490
-0.051
Minicheck-7B
0.980
0.020
0.500
0.000
0.773
0.010
0.392
-0.335
GLM-4-9B-chat
0.860
0.150
0.505
0.014
0.900
0.120
0.510
0.059
—CoT
0.800
0.240
0.520
0.048
0.750
0.140
0.445
-0.139
—Prompt-B
0.510
0.310
0.410
-0.184
0.500
0.430
0.465
-0.070
Llama3.1-8B-instruct
0.410
0.640
0.525
0.052
0.103
0.050
0.077
-0.837
—CoT
0.447
0.390
0.418
-0.164
0.060
0.023
0.042
-0.917
—Prompt-B
0.840
0.180
0.510
0.027
0.060
0.040
0.050
-0.900
—SFT
0.970
0.460
0.715
0.508
0.980
0.103
0.542
0.173
Qwen3-8B
0.973
0.163
0.568
0.233
0.803
0.413
0.608
0.237
—CoT
0.950
0.123
0.537
0.131
0.677
0.460
0.568
0.140
—Prompt-B
0.970
0.120
0.545
0.171
0.670
0.420
0.545
0.093
—SFT
0.973
0.573
0.773
0.596
0.910
0.630
0.770
0.563
Qwen3-14B
0.980
0.197
0.588
0.285
0.817
0.380
0.598
0.219
—CoT
1.000
0.190
0.595
0.324
0.830
0.317
0.573
0.171
—Prompt-B
0.950
0.170
0.560
0.192
0.680
0.390
0.535
0.073
Qwen3-32B
0.983
0.197
0.590
0.292
0.850
0.300
0.575
0.180
—CoT
0.980
0.210
0.595
0.298
0.823
0.330
0.577
0.176
—Prompt-B
0.850
0.290
0.570
0.169
0.870
0.330
0.600
0.238
—SFT
0.950
0.600
0.775
0.587
0.960
0.630
0.795
0.625
Commercial Models
Claude-4.5-Sonnet
0.760
0.750
0.755
0.510
0.610
0.720
0.665
0.332
—CoT
0.720
0.840
0.780
0.564
0.640
0.790
0.715
0.435
—Prompt-B
0.430
0.750
0.590
0.190
0.440
0.660
0.590
0.190
DeepSeek-v3
0.930
0.580
0.755
0.544
0.950
0.520
0.735
0.521
—CoT
0.930
0.540
0.735
0.510
0.930
0.440
0.685
0.424
Table 9: Detailed performance statistics in 64k, 100k scales. PromptB indicates that the target summary is positioned at the Beginning, while the default setup uses Prompt-E, which positions the target summary at the End. Voting Ensembles are abbreviation-coded as follows: (DSV3/DSR/C) denotes DeepSeek-V3, DeepSeek-R1, and Claude-sonnet-4-20250514-v1; (DSV4/GPT/C) denotes DeepSeek-V4, GPT-5.2-chat, and Claude-sonnet-4-20250514-v1.
Model
L (64k)
XL (100k)
Non-hallu
Hallu
BAcc.
MCC.
Non-hallu
Hallu
BAcc.
MCC.
Open-Source Models
InternLM2.5-20B
0.880
0.020
0.450
-0.196
0.960
0.000
0.480
-0.143
—CoT
0.960
0.040
0.500
0.000
0.980
0.000
0.490
-0.101
—Prompt-B
0.920
0.040
0.480
-0.084
0.780
0.080
0.430
-0.196
Minicheck-7B
0.820
0.020
0.420
-0.266
0.620
0.000
0.310
-0.484
GLM-4-9B-chat
1.000
0.000
0.500
0.000
1.000
0.000
0.500
0.000
—CoT
0.853
0.087
0.470
-0.094
0.980
0.040
0.510
0.059
—Prompt-B
0.500
0.300
0.400
-0.204
0.660
0.260
0.460
-0.087
Llama3.1-8B-instruct
0.000
0.000
-1.000
0.000
0.000
0.000
-1.000
0.000
—CoT
0.000
0.000
-1.000
0.000
0.000
0.000
-1.000
0.000
—Prompt-B
0.000
0.000
-1.000
0.000
0.000
0.000
-1.000
0.000
—SFT
0.260
0.160
0.210
-0.583
0.433
0.187
0.310
-0.583
Qwen3-8B
0.760
0.307
0.533
0.075
0.780
0.180
0.480
-0.050
—CoT
0.880
0.273
0.577
0.192
0.700
0.200
0.450
-0.115
—Prompt-B
0.840
0.380
0.610
0.248
0.680
0.500
0.590
0.183
—SFT
0.980
0.620
0.800
0.643
0.980
0.540
0.760
0.579
Qwen3-14B
0.773
0.320
0.547
0.105
0.700
0.333
0.517
0.036
—CoT
0.840
0.300
0.570
0.167
0.740
0.327
0.533
0.073
—Prompt-B
0.700
0.420
0.560
0.125
0.660
0.460
0.560
0.122
Qwen3-32B
0.920
0.273
0.597
0.253
0.840
0.320
0.580
0.187
—CoT
0.913
0.247
0.580
0.215
0.800
0.360
0.580
0.178
—Prompt-B
0.800
0.320
0.560
0.137
0.840
0.260
0.550
0.123
—SFT
0.960
0.540
0.750
0.551
0.940
0.500
0.720
0.490
Commercial Models
Claude-4.5-Sonnet
0.620
0.740
0.680
0.363
0.660
0.680
0.670
0.340
—CoT
0.640
0.700
0.670
0.341
0.700
0.800
0.750
0.503
—Prompt-B
0.240
0.560
0.400
-0.211
0.440
0.520
0.480
-0.040
DeepSeek-v3
0.980
0.380
0.680
0.450
0.980
0.360
0.670
0.433
—CoT
0.940
0.300
0.620
0.312
0.920
0.300
0.610
0.280
Table 10: Correspondence between hallucination types and prompts for generation.
Hallucination Type
Reference
Entity Hallucination
Fig. 8 Fig. 9
Numerical Hallucination
Fig. 10 Fig. 11
Relation Hallucination
Fig. 12 Fig. 13
Logical Inversion
Fig. 14 Fig. 15
Event Hallucination
Fig. 16 Fig. 17
Temporal Hallucination
Fig. 18 Fig. 19
Causal Hallucination
Fig. 20 Fig. 21
Event Fabrication
Fig. 22 Fig. 23
왜 중요한가
긴 문서를 요약해주는 AI 서비스가 늘고 있지만, 얼마나 정확한지 검증할 표준화된 시험지가 부족했다. 이 벤치마크는 실제 서비스에 쓰이는 모델들이 긴 글에서 어떤 유형의 실수를 얼마나 자주 저지르는지 구체적으로 보여줘, 더 신뢰할 수 있는 요약 AI를 만드는 데 기준점이 될 수 있다.
이 논문의 용어
환각(Hallucination) · AI가 원문에 없는 내용을 사실인 것처럼 지어내는 오류