Table 1: Comparison of our benchmark with other benchmarks in novel. ‘Summ. Halluc.’, ‘100K’, ‘Auto. Gen.’, and ‘Diff. Len.’ mean whether it is a summarization hallucination detection dataset, whether it reaches up to 100K tokens, whether the hallucinated data is generated through automated methods, and whether it encompasses different levels of length, respectively. LCHD Liu et al. (2025) refers to the long-context hallucination detection dataset.
Benchmark
Summ. Halluc.
100K
Auto. Label
Diff. Len.
Nocha (2024a)
✗
✓
✗
✗
StorySumm (2024)
✓
✗
✗
✗
FABLES (2024)
✓
✓
✗
✗
LCHD (2025)
✓
✗
✓
✗
CLIPPER (2025)
✗
✓
✓
✓
LongNovel (Ours)
✓
✓
✓
✓
Table 2: Statistics of the LongNovel training and test sets. Hallu. and Non-H. represent the counts of hallucinated and non-hallucinated samples, respectively.
Context
Train
Test
Length
Hallu.
Non-H.
Total
Hallu.
Non-H.
Total
S
1,454
1,430
2,884
100
100
200
M
1,292
1,178
2,470
100
100
200
L
–
–
–
50
50
100
XL
–
–
–
50
50
100
Total
5,354
600
Table 3: Token sequence length of LongNovel test set (values are in thousands, i.e., k).
Tokenizer
S (n=200)
M (n=200)
L (n=100)
XL (n=100)
Min
Mean
Max
Min
Mean
Max
Min
Mean
Max
Min
Mean
Max
GPT-4o
14.65
21.63
33.35
30.15
43.81
60.93
62.34
87.32
119.87
98.78
139.68
187.36
InternLM-2.5
13.37
15.86
18.01
29.31
32.21
34.35
59.62
64.03
66.98
92.81
100.68
104.17
Qwen-3
14.40
15.81
18.06
30.23
32.09
34.48
62.22
63.76
65.93
98.30
100.21
103.32
GLM-4
13.67
15.49
17.62
28.65
31.48
33.58
59.47
62.43
64.91
93.89
98.18
100.70
Table 4: Balanced Accuracy and MCC of various models. The best score is bold and the second-best is underline within each category (Open-Source vs Commercial). Voting Ensembles are abbreviation-coded as follows: (DSV3/DSR/C) denotes DeepSeek-V3, DeepSeek-R1, and Claude-sonnet-4-20250514-v1; (DSV4/GPT/C) denotes DeepSeek-V4, GPT-5.2-chat, and Claude-sonnet-4-20250514-v1.
Model
S(16k)
M(32k)
L(64k)
XL(100k)
BAcc.
MCC.
BAcc.
MCC.
BAcc.
MCC.
BAcc.
MCC.
Open-Source Models
InternLM2.5-20B
0.545
0.217
0.505
0.024
0.450
-0.196
0.480
-0.143
—CoT
0.555
0.241
0.520
0.084
0.500
0.000
0.490
-0.101
Minicheck-7B
0.500
0.000
0.392
-0.335
0.420
-0.266
0.310
-0.484
GLM-4-9B-chat
0.505
0.014
0.510
0.059
0.500
0.000
0.500
0.000
—CoT
0.520
0.048
0.445
-0.139
0.470
-0.094
0.510
0.059
Llama3.1-8B-instruct
0.525
0.052
0.077
-0.837
0.000
-1.000
0.000
-1.000
—CoT
0.418
-0.164
0.042
-0.917
0.000
-1.000
0.000
-1.000
—SFT
0.715
0.508
0.542
0.173
0.210
-0.583
0.310
-0.583
Qwen3-8B
0.568
0.233
0.608
0.237
0.533
0.075
0.480
-0.050
—CoT
0.537
0.131
0.568
0.140
0.577
0.192
0.450
-0.115
—SFT
0.773
0.596
0.770
0.563
0.800
0.643
0.760
0.579
Qwen3-14B
0.588
0.285
0.598
0.219
0.547
0.105
0.517
0.036
—CoT
0.595
0.324
0.573
0.171
0.570
0.167
0.533
0.073
Qwen3-32B
0.590
0.292
0.575
0.180
0.597
0.253
0.580
0.187
—CoT
0.595
0.298
0.577
0.176
0.580
0.215
0.580
0.178
—SFT
0.775
0.587
0.795
0.625
0.750
0.551
0.720
0.490
Commercial Models
Claude-4.5-Sonnet
0.755
0.510
0.665
0.332
0.680
0.363
0.670
0.340
—CoT
0.780
0.564
0.715
0.435
0.670
0.341
0.750
0.503
DeepSeek-v3
0.755
0.544
0.735
0.521
0.680
0.450
0.670
0.433
—CoT
0.735
0.510
0.685
0.424
0.620
0.312
0.610
0.280
—RAG
–
–
–
–
0.530
0.068
0.500
-0.014
DeepSeek-r1
0.775
0.592
0.735
0.510
0.640
0.382
0.680
0.450
—CoT
0.710
0.440
0.680
0.435
0.610
0.327
0.670
0.433
DeepSeek-v4-flash
0.785
0.573
0.810
0.625
0.830
0.660
0.830
0.660
GPT-5.2-chat
0.785
0.570
0.770
0.541
0.810
0.620
0.740
0.482
—CoT
0.775
0.555
0.765
0.530
0.780
0.560
0.750
0.503
—RAG
–
–
–
–
0.830
0.661
0.750
0.500
Table 5: Ablation study on data construction strategies across various models and dataset scales.
Model
S(16k)
M(32k)
L(64k)
XL(100k)
BAcc.
MCC.
BAcc.
MCC.
BAcc.
MCC.
BAcc.
MCC.
Qwen3-8B
0.568
0.233
0.608
0.237
0.533
0.075
0.480
-0.050
Qwen3-8B-MMA
0.577
0.248
0.577
0.171
0.610
0.261
0.710
0.447
Qwen3-8B-FULL
0.773
0.596
0.770
0.563
0.800
0.643
0.760
0.579
Llama3.1-8B-instruct
0.525
0.052
0.082
-0.837
0.000
-1.000
0.000
-1.000
Llama3.1-8B-MMA
0.585
0.290
0.502
0.008
0.020
-0.961
0.000
-1.000
Llama3.1-8B-FULL
0.720
0.508
0.542
0.173
0.210
-0.583
0.310
-0.583
Qwen3-32B
0.590
0.292
0.575
0.180
0.597
0.253
0.580
0.187
Qwen3-32B-MMA
0.605
0.343
0.587
0.281
0.610
0.352
0.680
0.469
Qwen3-32B-FULL
0.775
0.587
0.795
0.625
0.750
0.551
0.720
0.490
Table 6: Statistical distribution and composition of the LongNovel dataset across various subsets. ZH and EN denote Chinese and English data; Original denotes concatenated and compressed human-written summaries; ERHG denotes Entity-Referenced Hallucination Generation; MMA denotes Multi-Model Arbitration; Hallu. and Non-H. represent the counts of hallucinated and non-hallucinated samples
Subset
Total
ZH
EN
Original
ERHG
MMA
Hallu.
Non-H.
Train_16k
2884
1230
1654
814
963
1107
1454
1430
Train_32k
2470
1091
1379
734
937
799
1292
1178
Valid_16k
200
100
100
60
40
100
100
100
Valid_32k
200
100
100
60
40
100
100
100
Test_16k
200
100
100
64
48
88
100
100
Test_32k
200
100
100
69
61
70
100
100
Test_64k
100
50
50
41
27
32
50
50
Test_100k
100
54
46
38
20
42
50
50
Table 7: Detailed composition of the Chinese subset in the LongNovel dataset, listing the book titles and authors across the training, validation, and test splits.
Split
Book Title
Author
Training Set
Bright Eyes in the Dark
Er Dong Tu Zi
Destined
Mo Shu Bai
Misty Rain Tower
Yixi Yanyu
The Gentlemen of the City
Jin Shisichai
We Love Each Other So Much (Chinese edition)
Marcela Serrano
Little Confucian Immortal Seeking the Unknown
Youziyin
Turing’s Code
Feitian Yexiang
Folding Beijing
Hao Jingfang
I’m Waiting for You in the Memory
Xin Yiwu
Bu Yi Jian
Feng Diuzi
The Chronicle of Qingxi: Volume 1
Hong Zhuxia
The Chronicle of Qingxi: Volume 2
Hong Zhuxia
The Chronicle of Qingxi: Volume 3
Hong Zhuxia
The Space-Time Painter
Hai Ya
Once Gone
Qingshan Huangzhong
Validation Set
Yuhong
Ban Yu
Let Me Look at You
Xin Yiwu
I’m Waiting for You in the Memory
Xin Yiwu
My Garden
Teng Ping
Test Set
Bad Kids
Zijin Chen
Exclusive Possession
Ding Mo
Blood is Burning
Bainian Ruge
Everyone is a Protagonist Except Me
Cong Wen
The System Granted Me Longevity
Zi Ling Feng Xue
Mysterious Lotus Casebook
Teng Ping
Cicadas Sing the Setting Sun West
Yu Luo Zhu Leng
Ge Lu Ming: Volume 1
Qin Huai
Ge Lu Ming: Volume 2
Qin Huai
The Creatures That We Are
Peng Pai
Table 8: Detailed performance statistics in 16k, 32k scales. PromptB indicates that the target summary is positioned at the Beginning, while the default setup uses Prompt-E, which positions the target summary at the End. Voting Ensembles are abbreviation-coded as follows: (DSV3/DSR/C) denotes DeepSeek-V3, DeepSeek-R1, and Claude-sonnet-4-20250514-v1; (DSV4/GPT/C) denotes DeepSeek-V4, GPT-5.2-chat, and Claude-sonnet-4-20250514-v1.
Model
S(16k)
M(32k)
Non-H
Hallu
BAcc.
MCC.
Non-H
Hallu
BAcc.
MCC.
Open-Source Models
InternLM2.5-20B
1.000
0.090
0.545
0.217
0.960
0.050
0.505
0.024
—CoT
1.000
0.110
0.555
0.241
0.960
0.080
0.520
0.084
—Prompt-B
0.790
0.380
0.585
0.186
0.950
0.030
0.490
-0.051
Minicheck-7B
0.980
0.020
0.500
0.000
0.773
0.010
0.392
-0.335
GLM-4-9B-chat
0.860
0.150
0.505
0.014
0.900
0.120
0.510
0.059
—CoT
0.800
0.240
0.520
0.048
0.750
0.140
0.445
-0.139
—Prompt-B
0.510
0.310
0.410
-0.184
0.500
0.430
0.465
-0.070
Llama3.1-8B-instruct
0.410
0.640
0.525
0.052
0.103
0.050
0.077
-0.837
—CoT
0.447
0.390
0.418
-0.164
0.060
0.023
0.042
-0.917
—Prompt-B
0.840
0.180
0.510
0.027
0.060
0.040
0.050
-0.900
—SFT
0.970
0.460
0.715
0.508
0.980
0.103
0.542
0.173
Qwen3-8B
0.973
0.163
0.568
0.233
0.803
0.413
0.608
0.237
—CoT
0.950
0.123
0.537
0.131
0.677
0.460
0.568
0.140
—Prompt-B
0.970
0.120
0.545
0.171
0.670
0.420
0.545
0.093
—SFT
0.973
0.573
0.773
0.596
0.910
0.630
0.770
0.563
Qwen3-14B
0.980
0.197
0.588
0.285
0.817
0.380
0.598
0.219
—CoT
1.000
0.190
0.595
0.324
0.830
0.317
0.573
0.171
—Prompt-B
0.950
0.170
0.560
0.192
0.680
0.390
0.535
0.073
Qwen3-32B
0.983
0.197
0.590
0.292
0.850
0.300
0.575
0.180
—CoT
0.980
0.210
0.595
0.298
0.823
0.330
0.577
0.176
—Prompt-B
0.850
0.290
0.570
0.169
0.870
0.330
0.600
0.238
—SFT
0.950
0.600
0.775
0.587
0.960
0.630
0.795
0.625
Commercial Models
Claude-4.5-Sonnet
0.760
0.750
0.755
0.510
0.610
0.720
0.665
0.332
—CoT
0.720
0.840
0.780
0.564
0.640
0.790
0.715
0.435
—Prompt-B
0.430
0.750
0.590
0.190
0.440
0.660
0.590
0.190
DeepSeek-v3
0.930
0.580
0.755
0.544
0.950
0.520
0.735
0.521
—CoT
0.930
0.540
0.735
0.510
0.930
0.440
0.685
0.424
Table 9: Detailed performance statistics in 64k, 100k scales. PromptB indicates that the target summary is positioned at the Beginning, while the default setup uses Prompt-E, which positions the target summary at the End. Voting Ensembles are abbreviation-coded as follows: (DSV3/DSR/C) denotes DeepSeek-V3, DeepSeek-R1, and Claude-sonnet-4-20250514-v1; (DSV4/GPT/C) denotes DeepSeek-V4, GPT-5.2-chat, and Claude-sonnet-4-20250514-v1.
Model
L (64k)
XL (100k)
Non-hallu
Hallu
BAcc.
MCC.
Non-hallu
Hallu
BAcc.
MCC.
Open-Source Models
InternLM2.5-20B
0.880
0.020
0.450
-0.196
0.960
0.000
0.480
-0.143
—CoT
0.960
0.040
0.500
0.000
0.980
0.000
0.490
-0.101
—Prompt-B
0.920
0.040
0.480
-0.084
0.780
0.080
0.430
-0.196
Minicheck-7B
0.820
0.020
0.420
-0.266
0.620
0.000
0.310
-0.484
GLM-4-9B-chat
1.000
0.000
0.500
0.000
1.000
0.000
0.500
0.000
—CoT
0.853
0.087
0.470
-0.094
0.980
0.040
0.510
0.059
—Prompt-B
0.500
0.300
0.400
-0.204
0.660
0.260
0.460
-0.087
Llama3.1-8B-instruct
0.000
0.000
-1.000
0.000
0.000
0.000
-1.000
0.000
—CoT
0.000
0.000
-1.000
0.000
0.000
0.000
-1.000
0.000
—Prompt-B
0.000
0.000
-1.000
0.000
0.000
0.000
-1.000
0.000
—SFT
0.260
0.160
0.210
-0.583
0.433
0.187
0.310
-0.583
Qwen3-8B
0.760
0.307
0.533
0.075
0.780
0.180
0.480
-0.050
—CoT
0.880
0.273
0.577
0.192
0.700
0.200
0.450
-0.115
—Prompt-B
0.840
0.380
0.610
0.248
0.680
0.500
0.590
0.183
—SFT
0.980
0.620
0.800
0.643
0.980
0.540
0.760
0.579
Qwen3-14B
0.773
0.320
0.547
0.105
0.700
0.333
0.517
0.036
—CoT
0.840
0.300
0.570
0.167
0.740
0.327
0.533
0.073
—Prompt-B
0.700
0.420
0.560
0.125
0.660
0.460
0.560
0.122
Qwen3-32B
0.920
0.273
0.597
0.253
0.840
0.320
0.580
0.187
—CoT
0.913
0.247
0.580
0.215
0.800
0.360
0.580
0.178
—Prompt-B
0.800
0.320
0.560
0.137
0.840
0.260
0.550
0.123
—SFT
0.960
0.540
0.750
0.551
0.940
0.500
0.720
0.490
Commercial Models
Claude-4.5-Sonnet
0.620
0.740
0.680
0.363
0.660
0.680
0.670
0.340
—CoT
0.640
0.700
0.670
0.341
0.700
0.800
0.750
0.503
—Prompt-B
0.240
0.560
0.400
-0.211
0.440
0.520
0.480
-0.040
DeepSeek-v3
0.980
0.380
0.680
0.450
0.980
0.360
0.670
0.433
—CoT
0.940
0.300
0.620
0.312
0.920
0.300
0.610
0.280
Table 10: Correspondence between hallucination types and prompts for generation.
Although context windows have expanded significantly in recent years, hallucinations in long-context summarization remain a challenge. Long novels are better suited than news or papers for researching these hallucinations, due to their intrinsic information and detailed descriptions of events and dialogues. However, current research lacks a multi-scale benchmark for hallucination detection in long-context novel summarization and does not fully explore how hallucinations change as the context grows longer. In this study, we propose LongNovel, a multi-scale long-context bilingual (Chinese and English) novel benchmark for hallucination detection. This benchmark is constructed from 29 Chinese novels (ranging from 16k to 100k tokens) and chapter-level data from the BookSum dataset. We design 8 hallucination types and employ a combination of Multi-Model Arbitration and Entity-Referenced Hallucination Generation to ensure both data authenticity and a balanced distribution of hallucination categories. Furthermore, we manually revise the content in the test set to guarantee data reliability. Extensive experimental results demonstrate that LongNovel is a challenging benchmark. We release LongNovel for future research. https://github.com/BDML-lab/LongNovel
作者 · Ruizhi Zhang, Jinwei Chen, Xiangju Lu, He Yan, Mo Yu, Junmin Zhu, Wei Zhang