Table 1: Datasets used for training and evaluation, covering figurative language and cultural knowledge across Arabic varieties and regions. Cultural datasets (AraDiCE-Culture, ArabCulture, Palm) and figurative language datasets (FannOrFlop, Jawaher, Kinayat) are visually distinguished; ArabicMMLU serves as a control.
Poem–explanation pairs capturing poetic preference and aesthetic judgment across 14 genres, 12 eras
6,984
Arabic poetry
Train
Jawaher magdy-etal-2025-jawaher
Arabic proverb understanding and interpretation
800 (train), 198 (test)
20 Arabic varieties
Train, test
Kinayat attia-etal-2026-beyond
Egyptian Arabic idiom–explanation pairs
150
Egyptian Arabic
Test
ArabicMMLU koto-etal-2024-arabicmmlu
Arabic language and grammar MCQs (control)
980
Modern Standard Arabic
Train
Figure 9: Performance difference on different datasets for models fine-tuned on the ArabCulture dataset (diff = fine-tuned accuracy - base accuracy).
Table 2: Average evaluation results across three runs on Jawaher, Kinayat, and AraDiCE datasets. Results show accuracy scores (↑) for base models and models fine-tuned on different subsets.
Model
Jawaher
Kinayat
AraDiCE
Base
ALLaM-7B-Instruct
0.8990
0.8400
0.7537
Qwen3-8B
0.8131
0.6800
0.5037
Fanar-1-9B-Instruct
0.8906
0.7667
0.6870
Llama-3.1-8B-Instruct
0.6886
0.5956
0.5204
Fine-tuned on Palm Subset
ALLaM-7B-Instruct
0.8620
0.8333
0.7167
Qwen3-8B
0.8300
0.7178
0.5204
Fanar-1-9B-Instruct
0.8603
0.7467
0.6926
Llama-3.1-8B-Instruct
0.7205
0.6333
0.5481
Fine-tuned on ArabCulture subset
ALLaM-7B-Instruct
0.8906
0.8267
0.7500
Qwen3-8B
0.8148
0.7133
0.5185
Fanar-1-9B-Instruct
0.8805
0.7533
0.6870
Llama-3.1-8B-Instruct
0.7020
0.6200
0.4926
Fine-tuned on Jawaher
ALLaM-7B-Instruct
0.8990
0.8533
0.7407
Qwen3-8B
0.7896
0.6756
0.5130
Fanar-1-9B-Instruct
0.8906
0.7489
0.6907
Llama-3.1-8B-Instruct
0.7121
0.6778
0.5130
Fine-tuned on FannOrFlop
ALLaM-7B-Instruct
0.8889
0.8289
0.7352
Qwen3-8B
0.8333
0.7000
0.5111
Fanar-1-9B-Instruct
0.8973
0.7822
0.6944
Llama-3.1-8B-Instruct
0.7121
0.6644
0.5241
Fine-tuned on ArabicMMLU (control)
ALLaM-7B-Instruct
0.8906
0.7897
0.7833
Qwen3-8B
0.7795
0.6369
0.5389
Fanar-1-9B-Instruct
0.8838
0.7436
0.6926
Llama-3.1-8B-Instruct
0.7037
0.6164
0.5463
Figure 10: Performance difference on different datasets for models fine-tuned on the FannOrFlop dataset (diff = fine-tuned accuracy - base accuracy).
Table 3: Aggregate performance changes on culture by model and fine-tuning dataset.
Model
Train Set
Avg Δ %
Impr
Regr
Net
ALLaM
Jawaher
-1.30
21
28
-7
FannOrFlop
-1.87
15
25
-10
Fanar
Jawaher
+0.37
22
20
+2
FannOrFlop
+0.73
23
19
+4
Llama
Jawaher
-0.77
38
42
-4
FannOrFlop
+0.37
43
41
+2
Qwen
Jawaher
+0.87
10
5
+5
FannOrFlop
+0.70
13
9
+4
Figure 11: Performance difference on different datasets for models fine-tuned on the Jawaher dataset (diff = fine-tuned accuracy - base accuracy).
Table 4: Performance changes on cultural evaluation by country and topic category, aggregated across all models and fine-tuning configurations.
Impr
Regr
Net
By Country
Jordan
50
18
+32
Lebanon
30
29
+1
Qatar
26
28
−2
Palestine
30
35
−5
Egypt
22
34
−12
Syria
27
45
−18
By Topic
Food/Cuisine
18
3
+15
Traditional Games
18
5
+13
Other
66
65
+1
Holidays/Occasions
40
46
−6
History/Civilization
3
11
−8
Religion
2
10
−8
Traditional Clothing
17
31
−14
Figure 12: Performance difference on different datasets for models fine-tuned on the Palm dataset (diff = fine-tuned accuracy - base accuracy).
Table 5: Overall fine-tuning effect on idiom and proverb interpretation, aggregated across all models. Cultural fine-tuning uses ArabCulture and Palm; Poetry fine-tuning uses FannOrFlop.
Task
Impr
Regr
Net
Avg Δ%
Cultural Fine-tuning
Idioms
195
159
+36
+1.00
Proverbs
149
162
−13
−0.27
Poetry Fine-tuning
Idioms
100
58
+42
+2.33
Proverbs
97
73
+24
+1.01
Figure 13: Performance difference on different datasets for models fine-tuned on the ArabicMMLU baseline dataset (diff = fine-tuned accuracy - base accuracy).
Table 6: Aggregate fine-tuning results by model and fine-tuning dataset for both idiom and proverb interpretation. Cultural fine-tuning datasets (ArabCulture, Palm) and Poetry fine-tuning (FannOrFlop) are visually distinguished. Net refers to the total net improved predictions across all three seeds.
Idioms
Proverbs
Model
Train Set
Avg Δ%
Net
Avg Δ%
Net
ALLaM
ArabCulture
−1.33
−6
−0.84
−5
Palm
−0.67
−3
−3.70
−22
Poetry
−1.11
−5
−1.01
−6
Fanar
ArabCulture
−1.33
−6
−1.01
−6
Palm
−2.00
−9
−3.03
−18
Poetry
+1.56
+7
+0.67
+4
LLaMA
ArabCulture
+2.44
+11
+1.35
+8
Palm
+3.78
+17
+3.20
+19
Poetry
+6.89
+31
+2.36
+14
Qwen
ArabCulture
+3.33
+15
+0.17
+1
Palm
+3.78
+17
+1.68
+10
Poetry
+2.02
+9
+2.00
+12
Figure 14: Average performance difference between models fine-tuned on the full FannOrFlop dataset vs. FannOrFlop subset.
Table 7: Zero-shot evaluation results across three runs with different random seeds on Jawaher, Kinayat, and AraDiCE datasets. Results show accuracy scores (↑) for base models and models fine-tuned on different subsets.
Run 1 (seed=0)
Run 2 (seed=42)
Run 3 (seed=21)
Model
Jawaher
Kinayat
AraDiCE
Jawaher
Kinayat
AraDiCE
Jawaher
Kinayat
AraDiCE
Base
ALLaM-7B-Instruct
0.8889
0.8600
0.7667
0.9091
0.8267
0.7667
0.8990
0.8333
0.7278
Qwen3-8B
0.8030
0.6733
0.5278
0.8333
0.6867
0.5056
0.8030
0.6800
0.4778
Fanar-1-9B-Instruct
0.8788
0.7733
0.7222
0.9040
0.7267
0.7000
0.8889
0.8000
0.6389
Llama-3.1-8B-Instruct
0.7020
0.5933
0.5056
0.6919
0.5933
0.5167
0.6717
0.6000
0.5389
Fine-tuned on Palm Subset
ALLaM-7B-Instruct
0.8687
0.8533
0.7333
0.8737
0.8133
0.7111
0.8434
0.8333
0.7056
Qwen3-8B
0.8232
0.7000
0.5500
0.8535
0.7333
0.5111
0.8131
0.7200
0.5000
Fanar-1-9B-Instruct
0.8586
0.7467
0.7278
0.8687
0.7133
0.7000
0.8535
0.7800
0.6500
Llama-3.1-8B-Instruct
0.7475
0.6267
0.5000
0.7020
0.6267
0.5500
0.7121
0.6467
0.5944
Fine-tuned on ArabCulture subset
ALLaM-7B-Instruct
0.8889
0.8467
0.7778
0.9091
0.8133
0.7389
0.8737
0.8200
0.7333
Qwen3-8B
0.8131
0.7067
0.5778
0.8333
0.7200
0.5000
0.7980
0.7133
0.4778
Fanar-1-9B-Instruct
0.8838
0.7533
0.6944
0.8788
0.7200
0.7222
0.8788
0.7867
0.6444
Llama-3.1-8B-Instruct
0.7121
0.6200
0.4333
0.7071
0.5933
0.5000
0.6869
0.6467
0.5444
Fine-tuned on Jawaher
ALLaM-7B-Instruct
0.8990
0.8733
0.7778
0.9192
0.8333
0.7278
0.8788
0.8533
0.7167
Qwen3-8B
0.7727
0.6800
0.5444
0.8182
0.6800
0.5000
0.7778
0.6667
0.4944
Fanar-1-9B-Instruct
0.8889
0.7467
0.7278
0.8990
0.7200
0.7000
0.8838
0.7800
0.6444
Llama-3.1-8B-Instruct
0.7323
0.6733
0.5000
0.7172
0.6733
0.5056
0.6869
0.6867
0.5333
Fine-tuned on FannOrFlop
ALLaM-7B-Instruct
0.8838
0.8667
0.7778
0.9040
0.8133
0.7167
0.8788
0.8067
0.7111
Qwen3-8B
0.8182
0.7000
0.5444
0.8485
0.7067
0.5000
0.8333
0.6933
0.4889
Fanar-1-9B-Instruct
0.8939
0.7867
0.7222
0.8939
0.7533
0.7222
0.9040
0.8067
0.6389
Llama-3.1-8B-Instruct
0.7374
0.6733
0.5278
0.7020
0.6200
0.5056
0.6970
0.7000
0.5389
Fine-tuned on ArabicMMLU (control)
ALLaM-7B-Instruct
0.8990
0.8031
0.8000
0.8990
0.7846
0.7778
0.8737
0.7815
0.7722
Qwen3-8B
0.7677
0.6369
0.5611
0.7980
0.6031
0.5444
0.7727
0.6708
0.5111
Fanar-1-9B-Instruct
0.8838
0.7323
0.7167
0.8939
0.7415
0.7222
0.8737
0.7569
0.6389
Llama-3.1-8B-Instruct
0.7222
0.6000
0.5056
0.7172
0.6277
0.5500
0.6717
0.6215
0.5833
Figure 15: Average performance difference between models fine-tuned on the full ArabCulture dataset vs. ArabCulture subset.
Table 8: Zero-shot evaluation results across three runs with different random seeds on Jawaher, Kinayat, and AraDiCE datasets. Results show accuracy scores (↑) for models finetuned on FannOrFlop and ArabCulture full datasets.
Run 1 (seed=0)
Run 2 (seed=42)
Run 3 (seed=21)
Model
Jawaher
Kinayat
AraDiCE
Jawaher
Kinayat
AraDiCE
Jawaher
Kinayat
AraDiCE
Fine-tuned on FannOrFlop
ALLaM-7B-Instruct
0.8990
0.8600
0.7889
0.9343
0.8333
0.7167
0.8990
0.8467
0.7222
Qwen3-8B
0.8333
0.6867
0.5556
0.8586
0.7333
0.4944
0.8586
0.7200
0.5278
Fanar-1-9B-Instruct
0.8687
0.7533
0.7167
0.8636
0.7267
0.6889
0.8586
0.7867
0.6278
Llama-3.1-8B-Instruct
0.7424
0.5733
0.5167
0.6919
0.5467
0.4889
0.6970
0.6200
0.5556
Fine-tuned on ArabCulture
ALLaM-7B-Instruct
0.8990
0.8467
0.7722
0.9040
0.8333
0.7444
0.8838
0.8333
0.7389
Qwen3-8B
0.7879
0.6800
0.5500
0.8283
0.6800
0.5111
0.8030
0.6667
0.4889
Fanar-1-9B-Instruct
0.8737
0.7667
0.7056
0.8939
0.7200
0.7389
0.8838
0.7667
0.6611
Llama-3.1-8B-Instruct
0.6919
0.5867
0.4722
0.6515
0.5333
0.5222
0.6667
0.6133
0.5222
Table 9: Zero-shot evaluation results for ALLaM-7B-Instruct on Jawaher, Kinayat, and AraDiCE datasets. Results show accuracy scores (↑) comparing the default LoRA configuration against two ablation settings.
Configuration
Subset
Seed
Jawaher
Kinayat
AraDiCE
Default
Palm
0
0.8687
0.8533
0.7333
42
0.8737
0.8133
0.7111
21
0.8434
0.8333
0.7056
FannOrFlop
0
0.8838
0.8667
0.7778
42
0.9040
0.8133
0.7167
21
0.8788
0.8067
0.7111
r=16, α=32, lr=1e-5
Palm
0
0.8737
0.8467
0.7556
42
0.8838
0.8000
0.7278
21
0.8636
0.8200
0.7278
r=16, α=32, lr=5e-4
Palm
0
0.8131
0.8200
0.7389
42
0.8333
0.8267
0.7278
21
0.8081
0.8200
0.7611
FannOrFlop
0
0.8838
0.7933
0.7111
42
0.8838
0.7467
0.6889
21
0.8687
0.7867
0.6611
Table 10: Average zero-shot evaluation results for ALLaM-7B-Instruct across three runs on Jawaher, Kinayat, and AraDiCE datasets. Δ values indicate the difference relative to the default configuration for the corresponding finetuning dataset. Results show accuracy scores (↑).
Configuration
Dataset
Jawaher
ΔJawaher
Kinayat
ΔKinayat
AraDiCE
ΔAraDiCE
r=4, α=8, lr=5e-5
Palm Subset
0.8620
—
0.8333
—
0.7167
—
FannOrFlop
0.8889
—
0.8289
—
0.7352
—
r=16, α=32, lr=1e-5
Palm Subset
0.8737
+0.0118
0.8222
−0.0111
0.7370
+0.0204
r=16, α=32, lr=5e-4
Palm Subset
0.8182
−0.0438
0.8222
−0.0111
0.7426
+0.0259
FannOrFlop
0.8788
−0.0101
0.7756
−0.0533
0.6870
−0.0481
Table 11: Clustered aggregate effects, pooled across all four models (items resampled jointly). ICC is the intraclass correlation coefficient. Δ and CI bounds in percentage points. ∗ denotes p<0.05.
Task
Finetuned on
Δ
95% CI
p
ICC
Sig.
AraDiCE
Jawaher
−0.19
[−1.76,+1.39]
0.8101
0.018
AraDiCE
Poetry
+0.00
[−1.48,+1.48]
0.9940
−0.006
Kinayat
ArabCulture
+0.78
[−1.22,+2.72]
0.4353
−0.011
Kinayat
Palm
+1.22
[−1.00,+3.44]
0.2881
0.027
Kinayat
Poetry
+2.33
[+0.39,+4.28]
0.0207
0.009
∗
Jawaher
ArabCulture
−0.08
[−1.56,+1.39]
0.9195
0.078
Jawaher
Palm
−0.46
[−2.06,+1.18]
0.5815
0.040
Jawaher
Poetry
+1.01
[−0.63,+2.61]
0.2312
0.046
Table 12: Per-run results significant by exact McNemar’s test (p<0.05).
Model
Finetuned on
Task
Seed
Δ
95% CI
p
ALLaM
Poetry
AraDiCE
42
−5.00
[−8.89,−1.11]
0.0225
Qwen
Palm
Idioms
42
+4.67
[+1.33,+8.00]
0.0156
Llama
Poetry
Idioms
21
+10.00
[+3.33,+16.67]
0.0081
ALLaM
Palm
Proverbs
21
−5.56
[−9.60,−1.52]
0.0127
Fanar
Palm
Proverbs
42
−3.54
[−6.57,−1.01]
0.0391
Table 13: Top 10 most frequently improved and regressed questions from the AraDiCE dataset, with counts across all models fine-tuned on figurative language (proverbs and poetry).
Count
Country
Question
Top 10 Most Frequently Improved
9
Jordan
؟ندرألاب ناضمر راهنب راطفإلاب رهجلا ةبوقع وش
7
Jordan
؟ندرألاب جاوزلا ديلاقتو تاداع يه وش
6
Syria
؟ايروس لهأ اهيف لفتحيب يللا دايعألا يه وش
6
Jordan
؟ندرألاب لافطألل ةيبعشلا باعلألا يه وش
6
Jordan
؟ندرألاب ةروهشم راجشأو تاتابن رتكأ وش
5
Qatar
؟رطق يف ةروهشملا تايولحلا ونش
5
Palestine
؟نيطسلفب ةروهشملا تايولحلا يه وش
5
Qatar
؟ناصرقلا ةياور فلأ ونم
5
Lebanon
؟نانبلب ةيمسرلا لطعلا مايأ وش
4
Lebanon
؟نانبلب ناضمر راهنب راطفإلاب رهجلا ةبوقع وش
Top 10 Most Frequently Regressed
7
Egypt
؟رصم يف ةيمسرلا لطعلا مايأ هيإ
7
Syria
؟لاجرلل يروسلا يديلقتلا سبللا وش
7
Syria
؟ناوسنلل يروسلا يديلقتلا سبللا وش
6
Palestine
؟نيطسلفب تراص ةيخيرات كراعم ٣ رهشأ وش
6
Lebanon
نانبلب نوجسلا نع يكحتب نيتياور نيوانع يدب
5
Syria
؟قشمد حتف داق يللا يباحصلا نيم
5
Syria
؟ايروس اهيلع لطتب يللا راحبلا يه وش
4
Qatar
؟رطق يف لالقتسالا ديع ىتم
4
Egypt
؟رصم يف جاوزلا ديلاقتو تاداع هيإ
4
Qatar
؟رطق يف ةماعلا تالصاوملل نيتليسو رهشأ ونش
Table 14: Representative subset of the 21 unstable idioms (improved in some seeds, regressed in others), with counts across all fine-tuned models.
Idiom (Arabic)
Impr
Regr
ْهشِو ْلَكَأ
9
3
ْهَدْلِجْلا ىَلَع
8
4
نِرِي ْهاَّلَخ
2
7
ْهَضاَخَمْلا ِّسَج
2
5
ْبلْقِتِو ْبَرْضِتِب اَيْنُّدلا
4
5
بياس هفك
6
2
ْرْحَبْلا ِتْحَف
4
1
Table 15: Representative subset of the 15 unstable proverbs (improved in some seeds, regressed in others).
Proverb (Arabic)
Dialect
Impr
Regr
لمج عسوي بابلا
Kuwaiti
1
8
اهنب نم هسفن لماع
Egyptian
2
7
اودعا بيجت ام نطبلا
Libyan
1
6
صوردلا وارعني ام ژاژغتلا تعاس
Mauritanian
5
3
مادك نم ديأو هرو نم ديأ عجر
Iraqi
4
2
اشاب نايمعلا ىلع روعألا
Omani
1
2
Table 16: Representative subset of the 15 unstable idioms (improved in some seeds, regressed in others) under poetry fine-tuning, with counts across all models.
Idiom (Arabic)
Impr
Regr
ْهشِو ْلَكَأ
3
3
نِرِي ْهاَّلَخ
1
4
ْهَدْلِجْلا ىَلَع
4
2
ْبلْقِتِو ْبَرْضِتِب اَيْنُّدلا
4
1
ْهَضاَخَمْلا ِّسَج
1
2
يِدْنه ْرْمَت ْنَبَل ْكَمَس
1
2
اهاَطَغ ّْدَرْو هَمْلِك
1
2
بياس هفك
2
1
شِو ْشوُلاَم
2
1
Table 17: The 7 unstable proverbs (improved in some seeds, regressed in others) under poetry fine-tuning.
Proverb (Arabic)
Dialect
Impr
Regr
حازم ةودعلا
Algerian
3
2
اهنب نم هسفن لماع
Egyptian
1
3
لبج لوقت تناو لمج لوقا انا
Omani
1
1
نيَعلا يحَتسِت مَفلا مَعطِا
Palestinian
2
1
يراوشلا زنخت ةدحو ةتوح
Moroccan
2
1
مادك نم ديأو هرو نم ديأ عجر
Iraqi
2
1
اهصنب ىضري ةزببخ ىضر ام يللا
Algerian
2
1
Table 18: Net fine-tuning effect by Arabic dialect variety on the proverb task.
Dialect
Impr
Regr
Net
Mauritanian
30
10
+20
Yemeni
15
1
+14
Iraqi
7
2
+5
Tunisian
7
2
+5
Jordanian
11
7
+4
Kuwaiti
12
8
+4
MSA
4
0
+4
Moroccan
7
6
+1
Qatari
4
3
+1
Egyptian
8
13
−5
Syrian
3
8
−5
Lebanese
7
12
−5
Palestinian
4
10
−6
Saudi
6
13
−7
Omani
3
10
−7
Libyan
5
13
−8
Sudanese
6
16
−10
Algerian
10
21
−11
Table 19: Improvement and regression breakdown on Kinayat idioms across cultural fine-tuning (ArabCulture and Palm) and poetry fine-tuning (FannOrFlop). Base% and FT% are accuracy before and after fine-tuning; Δ% is the percentage-point change, reported with 95% paired-bootstrap confidence intervals over test items and exact McNemar p-values; Impr and Regr are the number of individual predictions improved or worsened. ∗ marks p<0.05.
Model / Dataset
Base%
FT%
Δ%
95% CI
p
Impr
Regr
ALLaM
ArabCulture (seed 0)
86.0
84.7
−1.33
[−5.33,+2.67]
0.754
4
6
ArabCulture (seed 21)
83.3
82.0
−1.33
[−6.00,+3.33]
0.774
5
7
ArabCulture (seed 42)
82.7
81.3
−1.33
[−5.33,+2.67]
0.754
4
6
Palm (seed 0)
86.0
85.3
−0.67
[−5.33,+4.00]
1.000
6
7
Palm (seed 21)
83.3
83.3
±0.00
[−4.67,+4.67]
1.000
6
6
Palm (seed 42)
82.7
81.3
−1.33
[−6.67,+4.00]
0.804
7
9
FannOrFlop (seed 0)
86.0
86.7
+0.67
[−3.33,+4.67]
1.000
5
4
FannOrFlop (seed 21)
83.3
80.7
−2.67
[−7.33,+2.00]
0.388
4
8
FannOrFlop (seed 42)
82.7
81.3
−1.33
[−5.33,+2.67]
0.754
4
6
Fanar
ArabCulture (seed 0)
77.3
75.3
−2.00
[−5.33,+1.33]
0.453
2
5
ArabCulture (seed 21)
80.0
78.7
−1.33
[−4.67,+2.00]
0.688
2
4
ArabCulture (seed 42)
72.7
72.0
−0.67
[−4.00,+2.67]
1.000
3
4
Palm (seed 0)
77.3
74.7
−2.67
[−6.67,+0.67]
0.289
2
6
Palm (seed 21)
80.0
78.0
−2.00
[−5.33,+0.67]
0.375
1
4
Palm (seed 42)
72.7
71.3
−1.33
[−4.67,+2.00]
0.688
2
4
FannOrFlop (seed 0)
77.3
78.7
+1.33
[−2.00,+4.67]
0.688
4
2
FannOrFlop (seed 21)
80.0
80.7
+0.67
[−2.00,+3.33]
1.000
3
2
FannOrFlop (seed 42)
72.7
75.3
+2.67
[−0.67,+6.67]
0.289
6
2
LLaMA
ArabCulture (seed 0)
59.3
62.0
+2.67
[−5.33,+10.67]
0.627
21
17
ArabCulture (seed 21)
60.0
64.7
+4.67
[−2.00,+11.33]
0.248
17
10
ArabCulture (seed 42)
59.3
59.3
±0.00
[−7.33,+6.67]
1.000
14
14
Palm (seed 0)
59.3
62.7
+3.33
[−4.67,+12.00]
0.533
23
18
Palm (seed 21)
60.0
64.7
+4.67
[−2.67,+12.67]
0.310
21
14
Palm (seed 42)
59.3
62.7
+3.33
[−4.00,+10.67]
0.487
19
14
FannOrFlop (seed 0)
59.3
67.3
+8.00
[+0.00,+16.00]
0.073
25
13
FannOrFlop (seed 21)
60.0
70.0
+10.00∗
[+3.33,+16.67]
0.008
22
7
FannOrFlop (seed 42)
59.3
62.0
+2.67
[−4.00,+9.33]
0.557
15
11
Table 20: Improvement and regression breakdown on Jawaher proverbs across cultural fine-tuning (ArabCulture and Palm) and poetry fine-tuning (FannOrFlop). Base% and FT% are accuracy before and after fine-tuning; Δ% is the percentage-point change, reported with 95% paired-bootstrap confidence intervals over test items and exact McNemar p-values; Impr and Regr are the number of individual predictions improved or worsened. ∗ marks p<0.05.
Model / Dataset
Base%
FT%
Δ%
95% CI
p
Impr
Regr
ALLaM
ArabCulture (seed 0)
88.9
88.9
±0.00
[−2.53,+2.53]
1.000
3
3
ArabCulture (seed 21)
89.9
87.4
−2.53
[−5.05,+0.00]
0.125
1
6
ArabCulture (seed 42)
90.9
90.9
±0.00
[−3.03,+3.03]
1.000
5
5
Palm (seed 0)
88.9
86.9
−2.02
[−5.56,+1.52]
0.388
4
8
Palm (seed 21)
89.9
84.3
−5.56∗
[−9.60,−1.52]
0.013
3
14
Palm (seed 42)
90.9
87.4
−3.54
[−7.58,+0.51]
0.143
5
12
FannOrFlop (seed 0)
88.9
88.4
−0.51
[−3.54,+2.53]
1.000
4
5
FannOrFlop (seed 21)
89.9
87.9
−2.02
[−5.05,+0.51]
0.289
2
6
FannOrFlop (seed 42)
90.9
90.4
−0.51
[−4.04,+3.03]
1.000
5
6
Fanar
ArabCulture (seed 0)
87.9
88.4
+0.51
[−2.02,+3.03]
1.000
4
3
ArabCulture (seed 21)
88.9
87.9
−1.01
[−4.04,+2.02]
0.754
4
6
ArabCulture (seed 42)
90.4
87.9
−2.53
[−6.06,+0.51]
0.227
3
8
Palm (seed 0)
87.9
85.9
−2.02
[−5.56,+1.52]
0.388
4
8
Palm (seed 21)
88.9
85.4
−3.54
[−7.07,−0.51]
0.065
2
9
Palm (seed 42)
90.4
86.9
−3.54∗
[−6.57,−1.01]
0.039
1
8
FannOrFlop (seed 0)
87.9
89.4
+1.52
[−2.02,+5.05]
0.581
8
5
FannOrFlop (seed 21)
88.9
90.4
+1.52
[−1.52,+5.05]
0.549
7
4
FannOrFlop (seed 42)
90.4
89.4
−1.01
[−4.55,+2.53]
0.774
5
7
LLaMA
ArabCulture (seed 0)
70.2
71.2
+1.01
[−3.54,+5.56]
0.824
11
9
ArabCulture (seed 21)
67.2
68.7
+1.52
[−3.03,+6.06]
0.678
13
10
ArabCulture (seed 42)
69.2
70.7
+1.52
[−3.54,+6.57]
0.690
14
11
Palm (seed 0)
70.2
74.7
+4.55
[+0.00,+9.09]
0.078
15
6
Palm (seed 21)
67.2
71.2
+4.04
[−0.51,+9.09]
0.152
16
8
Palm (seed 42)
69.2
70.2
+1.01
[−3.54,+5.56]
0.832
12
10
FannOrFlop (seed 0)
70.2
73.7
+3.54
[−1.01,+8.08]
0.210
15
8
FannOrFlop (seed 21)
67.2
69.7
+2.53
[−2.53,+7.58]
0.442
16
11
FannOrFlop (seed 42)
69.2
70.2
+1.01
[−4.04,+6.06]
0.845
14
12
Table 21: Improvement and regression breakdown on AraDiCE-Culture across figurative fine-tuning (Jawaher) and poetry fine-tuning (FannOrFlop). Δ% is the percentage-point change from base to fine-tuned accuracy, reported with 95% paired-bootstrap confidence intervals over test items; Impr and Regr are the number of individual predictions improved or worsened. ∗ marks p<0.05 under an exact McNemar test.
Model / Dataset
Δ%
95% CI
Impr
Regr
ALLaM
Jawaher (seed 0)
+1.11
[−3.33,+5.56]
10
8
Jawaher (seed 21)
−1.11
[−5.56,+3.33]
8
10
Jawaher (seed 42)
−3.89
[−7.78,+0.00]
3
10
FannOrFlop (seed 0)
+1.11
[−3.33,+5.56]
9
7
FannOrFlop (seed 21)
−1.67
[−5.56,+1.67]
4
7
FannOrFlop (seed 42)
−5.00∗
[−8.89,−1.11]
2
11
Fanar
Jawaher (seed 0)
+0.56
[−3.89,+5.00]
9
8
Jawaher (seed 21)
+0.56
[−3.89,+5.00]
8
7
Jawaher (seed 42)
±0.00
[−3.33,+3.33]
5
5
FannOrFlop (seed 0)
±0.00
[−3.89,+3.89]
7
7
FannOrFlop (seed 21)
±0.00
[−4.44,+4.44]
8
8
FannOrFlop (seed 42)
+2.22
[−1.67,+6.11]
8
4
LLaMA
Jawaher (seed 0)
−0.56
[−6.11,+5.00]
12
13
Jawaher (seed 21)
−0.56
[−5.56,+4.44]
11
12
Jawaher (seed 42)
−1.11
[−7.22,+5.00]
15
17
FannOrFlop (seed 0)
+2.22
[−3.33,+7.78]
16
12
FannOrFlop (seed 21)
±0.00
[−5.56,+5.56]
13
13
FannOrFlop (seed 42)
−1.11
[−7.22,+5.00]
14
16
Qwen
Jawaher (seed 0)
+1.67
[−1.11,+4.44]
5
2
Jawaher (seed 21)
+1.67
[+0.00,+3.89]
3
0
Jawaher (seed 42)
−0.56
[−2.78,+1.67]
2
3
FannOrFlop (seed 0)
+1.67
[−1.67,+5.00]
6
3
FannOrFlop (seed 21)
+1.11
[−1.67,+3.89]
4
2
FannOrFlop (seed 42)
−0.56
[−3.33,+2.22]
3
4
Table 22: Top 15 most frequently improved and regressed idioms across all models fine-tuned on cultural data (ArabCulture and Palm).
Count
Idiom
Top 15 Most Frequently Improved
16
ْدوُع ىَلَع ْدوُد
9
ْهَديِدَحْلا ىَلَع
9
ْهشِو ْلَكَأ
8
ْهَدْلِجْلا ىَلَع
6
ْعاَرِّدلاو ْعاَبْلاِب
6
ْماَّدُق ْنِم ْديإِو اَرَو ْنِم ْديإ
6
بياس هفك
6
ْهَفْشاَن ْهُديإ
6
ّْبَد ْنِمْو ِّبَه ْنِم
6
ْتيبلا ِنِم ْهَريِمَخْلا ِعَطْقِي ْهُّشِو
6
ْداَّدَحْلا ِعَنَص اَم ْهُنيبْو هُنيب
6
ْروُّزلا ِنِم ْشْلِزْنِي اَم
5
ْهُنيِع ْنِم ْعِلِط
4
ْقِطاَّنلا ْقِلاَخْلا
4
ىَراَصَن ْةِزاَوَج
Top 15 Most Frequently Regressed
9
ْساَّنلا ِّيَز
7
نِرِي ْهاَّلَخ
6
ْهُبوُت ْنِم ْشوُم
6
ْهُغاَد ْباَج
6
ْلاَخْلُخْلاِب اَيْنُّدلا هُدْنَع
6
ْلِجْنِمْلاِو ْلِجْنِحْلاِب
6
ْهُتِّبُق يِف اَهْباَج
6
يفيِص ْهَخيِّطَب ْهُنْطَب يف طَح
6
رَبْلا اَهْبِياَج ْشوُم
5
ْبلْقِتِو ْبَرْضِتِب اَيْنُّدلا
5
هَنيِحْط ِرْحَبلا ِلَمَع
5
ْهَضاَخَمْلا ِّسَج
5
ْهُعاَبُص ِّلِض يِف ىَراَّدإ
Table 23: Top 15 most frequently improved and regressed proverbs across all fine-tuned models, with counts and dialect variety.
Figurative language is deeply culturally embedded; fluent use requires not just linguistic competence but cultural immersion. We ask whether LLMs can learn this link: does fine-tuning on cultural data improve figurative language understanding, and vice versa? We conduct a systematic study across four models (ALLaM-7B, Fanar-1-9B, Qwen3-8B, Llama-3.1-8B) and six Arabic datasets spanning cultural commonsense, proverbs, and poetry across diverse dialects and regions. Fine-tuning on poetry improves idiom comprehension (+2.33%, p<0.05), a gain our ArabicMMLU control does not reproduce, indicating that it stems from figurative content rather than Arabic language adaptation and pointing to a sensitivity to non-literal meaning that transfers across figurative types. Cultural fine-tuning, by contrast, lowers proverb-interpretation accuracy in both Arabic-centric models. Transfer between the two domains is otherwise indistinguishable from noise, with Arabic models frequently regressing after fine-tuning, suggesting prior saturation of relevant knowledge, while multilingual models show greater adaptation headroom. Error analysis further reveals that fine-tuning reinforces experiential cultural knowledge while destabilizing historically grounded factual knowledge. Our findings suggest that the relationship between culture and figurative language, though conceptually natural, is not straightforwardly captured through fine-tuning alone.