Figurative and Cultural Knowledge in LLMs: Investigating Cross-Domain Transfer through Fine-Tuning
속담·시·문화 지식을 AI에 따로 학습시켜도 서로 잘 옮겨붙지 않는다
연구진은 아랍어 모델 4종에 문화 상식 데이터와 속담·시 같은 비유적 언어 데이터를 각각 추가 학습시켜, 한쪽 지식이 다른 쪽 이해력을 끌어올리는지 확인했다. 결과적으로 시를 학습시키면 관용구 이해력이 base 모델 대비 2.33% 올랐지만, 이는 통계적으로 유의미한 유일한 효과였고 나머지 조합은 대부분 잡음 수준이거나 오히려 성능이 떨어졌다. 문화 데이터를 학습시키면 오히려 속담 해석 정확도가 떨어지는 경우도 있었다.
METAL MEDIA 해설 도표
속담·시·문화 지식을 AI에 따로 학습시켜도 서로 잘 옮겨붙지 않는다
01ALLaM-7B, Fanar-1-9B, Qwen3-8B, Llama-3.1-8B 네 모델에 아랍어 문화 상식 데이터셋(ArabCulture, Palm)과 비유 언어 데이터셋(FannOrFlop 시, Jawaher 속담)을 LoRA 방식으로 각각 미세조정한 뒤, 기존 모델(base)과 성능 차이를 비교했다
02단순히 아랍어 텍스트에 노출된 효과인지 구분하기 위해 ArabicMMLU를 대조군으로 별도 학습시켰다
03시(FannOrFlop) 데이터로 학습한 모델은 관용구 벤치마크 Kinayat에서 평균 2.33% 향상됐고, 이는 신뢰구간이 0을 벗어나는 유일한 결과였다(p=0.021), 반면 대조군인 ArabicMMLU 학습은 같은 벤치마크에서 오히려 성능이 떨어져 이 효과가 아랍어 적응이 아닌 비유적 내용 자체 때문임을 보여준다
04문화 데이터(Palm)로 학습하면 ALLaM-7B와 Fanar-1-9B 두 아랍어 특화 모델의 속담 이해력이 각각 3.70%, 3.03% 떨어졌는데, 이 두 결과가 이 연구에서 통계적으로 뒷받침되는 유일한 부정적 효과였다
05아랍어 특화 모델들은 이미 사전학습으로 관련 지식을 많이 흡수한 상태라 추가 학습이 오히려 방해가 되는 경향을 보였고, 다국어 모델(Qwen3, Llama-3.1)은 더 여지 있게 반응했다
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.
무엇을 했나
ALLaM-7B, Fanar-1-9B, Qwen3-8B, Llama-3.1-8B 네 모델에 아랍어 문화 상식 데이터셋(ArabCulture, Palm)과 비유 언어 데이터셋(FannOrFlop 시, Jawaher 속담)을 LoRA 방식으로 각각 미세조정한 뒤, 기존 모델(base)과 성능 차이를 비교했다
단순히 아랍어 텍스트에 노출된 효과인지 구분하기 위해 ArabicMMLU를 대조군으로 별도 학습시켰다
시(FannOrFlop) 데이터로 학습한 모델은 관용구 벤치마크 Kinayat에서 평균 2.33% 향상됐고, 이는 신뢰구간이 0을 벗어나는 유일한 결과였다(p=0.021), 반면 대조군인 ArabicMMLU 학습은 같은 벤치마크에서 오히려 성능이 떨어져 이 효과가 아랍어 적응이 아닌 비유적 내용 자체 때문임을 보여준다
문화 데이터(Palm)로 학습하면 ALLaM-7B와 Fanar-1-9B 두 아랍어 특화 모델의 속담 이해력이 각각 3.70%, 3.03% 떨어졌는데, 이 두 결과가 이 연구에서 통계적으로 뒷받침되는 유일한 부정적 효과였다
아랍어 특화 모델들은 이미 사전학습으로 관련 지식을 많이 흡수한 상태라 추가 학습이 오히려 방해가 되는 경향을 보였고, 다국어 모델(Qwen3, Llama-3.1)은 더 여지 있게 반응했다
오류 분석 결과 미세조정은 명절·음식·전통놀이 같은 경험적 문화 지식은 강화하지만, 역사·정치 등 사실 기반 지식은 오히려 불안정하게 만드는 경향을 보였다
Figure 2: Average performance difference.
Table 1: Datasets used for training and evaluation, covering figurative language and cultural knowledge across Arabic varieties and regions. Cultural datasets (AraDiCE-Culture, ArabCulture, Palm) and figurative language datasets (FannOrFlop, Jawaher, Kinayat) are visually distinguished; ArabicMMLU serves as a control.
Poem–explanation pairs capturing poetic preference and aesthetic judgment across 14 genres, 12 eras
6,984
Arabic poetry
Train
Jawaher magdy-etal-2025-jawaher
Arabic proverb understanding and interpretation
800 (train), 198 (test)
20 Arabic varieties
Train, test
Kinayat attia-etal-2026-beyond
Egyptian Arabic idiom–explanation pairs
150
Egyptian Arabic
Test
ArabicMMLU koto-etal-2024-arabicmmlu
Arabic language and grammar MCQs (control)
980
Modern Standard Arabic
Train
Figure 9: Performance difference on different datasets for models fine-tuned on the ArabCulture dataset (diff = fine-tuned accuracy - base accuracy).
Table 2: Average evaluation results across three runs on Jawaher, Kinayat, and AraDiCE datasets. Results show accuracy scores (↑) for base models and models fine-tuned on different subsets.
Model
Jawaher
Kinayat
AraDiCE
Base
ALLaM-7B-Instruct
0.8990
0.8400
0.7537
Qwen3-8B
0.8131
0.6800
0.5037
Fanar-1-9B-Instruct
0.8906
0.7667
0.6870
Llama-3.1-8B-Instruct
0.6886
0.5956
0.5204
Fine-tuned on Palm Subset
ALLaM-7B-Instruct
0.8620
0.8333
0.7167
Qwen3-8B
0.8300
0.7178
0.5204
Fanar-1-9B-Instruct
0.8603
0.7467
0.6926
Llama-3.1-8B-Instruct
0.7205
0.6333
0.5481
Fine-tuned on ArabCulture subset
ALLaM-7B-Instruct
0.8906
0.8267
0.7500
Qwen3-8B
0.8148
0.7133
0.5185
Fanar-1-9B-Instruct
0.8805
0.7533
0.6870
Llama-3.1-8B-Instruct
0.7020
0.6200
0.4926
Fine-tuned on Jawaher
ALLaM-7B-Instruct
0.8990
0.8533
0.7407
Qwen3-8B
0.7896
0.6756
0.5130
Fanar-1-9B-Instruct
0.8906
0.7489
0.6907
Llama-3.1-8B-Instruct
0.7121
0.6778
0.5130
Fine-tuned on FannOrFlop
ALLaM-7B-Instruct
0.8889
0.8289
0.7352
Qwen3-8B
0.8333
0.7000
0.5111
Fanar-1-9B-Instruct
0.8973
0.7822
0.6944
Llama-3.1-8B-Instruct
0.7121
0.6644
0.5241
Fine-tuned on ArabicMMLU (control)
ALLaM-7B-Instruct
0.8906
0.7897
0.7833
Qwen3-8B
0.7795
0.6369
0.5389
Fanar-1-9B-Instruct
0.8838
0.7436
0.6926
Llama-3.1-8B-Instruct
0.7037
0.6164
0.5463
Figure 10: Performance difference on different datasets for models fine-tuned on the FannOrFlop dataset (diff = fine-tuned accuracy - base accuracy).
Table 3: Aggregate performance changes on culture by model and fine-tuning dataset.
Model
Train Set
Avg Δ %
Impr
Regr
Net
ALLaM
Jawaher
-1.30
21
28
-7
FannOrFlop
-1.87
15
25
-10
Fanar
Jawaher
+0.37
22
20
+2
FannOrFlop
+0.73
23
19
+4
Llama
Jawaher
-0.77
38
42
-4
FannOrFlop
+0.37
43
41
+2
Qwen
Jawaher
+0.87
10
5
+5
FannOrFlop
+0.70
13
9
+4
Figure 11: Performance difference on different datasets for models fine-tuned on the Jawaher dataset (diff = fine-tuned accuracy - base accuracy).
Table 4: Performance changes on cultural evaluation by country and topic category, aggregated across all models and fine-tuning configurations.
Impr
Regr
Net
By Country
Jordan
50
18
+32
Lebanon
30
29
+1
Qatar
26
28
−2
Palestine
30
35
−5
Egypt
22
34
−12
Syria
27
45
−18
By Topic
Food/Cuisine
18
3
+15
Traditional Games
18
5
+13
Other
66
65
+1
Holidays/Occasions
40
46
−6
History/Civilization
3
11
−8
Religion
2
10
−8
Traditional Clothing
17
31
−14
Figure 12: Performance difference on different datasets for models fine-tuned on the Palm dataset (diff = fine-tuned accuracy - base accuracy).
Table 5: Overall fine-tuning effect on idiom and proverb interpretation, aggregated across all models. Cultural fine-tuning uses ArabCulture and Palm; Poetry fine-tuning uses FannOrFlop.
Task
Impr
Regr
Net
Avg Δ%
Cultural Fine-tuning
Idioms
195
159
+36
+1.00
Proverbs
149
162
−13
−0.27
Poetry Fine-tuning
Idioms
100
58
+42
+2.33
Proverbs
97
73
+24
+1.01
Figure 13: Performance difference on different datasets for models fine-tuned on the ArabicMMLU baseline dataset (diff = fine-tuned accuracy - base accuracy).
Table 6: Aggregate fine-tuning results by model and fine-tuning dataset for both idiom and proverb interpretation. Cultural fine-tuning datasets (ArabCulture, Palm) and Poetry fine-tuning (FannOrFlop) are visually distinguished. Net refers to the total net improved predictions across all three seeds.
Idioms
Proverbs
Model
Train Set
Avg Δ%
Net
Avg Δ%
Net
ALLaM
ArabCulture
−1.33
−6
−0.84
−5
Palm
−0.67
−3
−3.70
−22
Poetry
−1.11
−5
−1.01
−6
Fanar
ArabCulture
−1.33
−6
−1.01
−6
Palm
−2.00
−9
−3.03
−18
Poetry
+1.56
+7
+0.67
+4
LLaMA
ArabCulture
+2.44
+11
+1.35
+8
Palm
+3.78
+17
+3.20
+19
Poetry
+6.89
+31
+2.36
+14
Qwen
ArabCulture
+3.33
+15
+0.17
+1
Palm
+3.78
+17
+1.68
+10
Poetry
+2.02
+9
+2.00
+12
Figure 14: Average performance difference between models fine-tuned on the full FannOrFlop dataset vs. FannOrFlop subset.
Table 7: Zero-shot evaluation results across three runs with different random seeds on Jawaher, Kinayat, and AraDiCE datasets. Results show accuracy scores (↑) for base models and models fine-tuned on different subsets.
Run 1 (seed=0)
Run 2 (seed=42)
Run 3 (seed=21)
Model
Jawaher
Kinayat
AraDiCE
Jawaher
Kinayat
AraDiCE
Jawaher
Kinayat
AraDiCE
Base
ALLaM-7B-Instruct
0.8889
0.8600
0.7667
0.9091
0.8267
0.7667
0.8990
0.8333
0.7278
Qwen3-8B
0.8030
0.6733
0.5278
0.8333
0.6867
0.5056
0.8030
0.6800
0.4778
Fanar-1-9B-Instruct
0.8788
0.7733
0.7222
0.9040
0.7267
0.7000
0.8889
0.8000
0.6389
Llama-3.1-8B-Instruct
0.7020
0.5933
0.5056
0.6919
0.5933
0.5167
0.6717
0.6000
0.5389
Fine-tuned on Palm Subset
ALLaM-7B-Instruct
0.8687
0.8533
0.7333
0.8737
0.8133
0.7111
0.8434
0.8333
0.7056
Qwen3-8B
0.8232
0.7000
0.5500
0.8535
0.7333
0.5111
0.8131
0.7200
0.5000
Fanar-1-9B-Instruct
0.8586
0.7467
0.7278
0.8687
0.7133
0.7000
0.8535
0.7800
0.6500
Llama-3.1-8B-Instruct
0.7475
0.6267
0.5000
0.7020
0.6267
0.5500
0.7121
0.6467
0.5944
Fine-tuned on ArabCulture subset
ALLaM-7B-Instruct
0.8889
0.8467
0.7778
0.9091
0.8133
0.7389
0.8737
0.8200
0.7333
Qwen3-8B
0.8131
0.7067
0.5778
0.8333
0.7200
0.5000
0.7980
0.7133
0.4778
Fanar-1-9B-Instruct
0.8838
0.7533
0.6944
0.8788
0.7200
0.7222
0.8788
0.7867
0.6444
Llama-3.1-8B-Instruct
0.7121
0.6200
0.4333
0.7071
0.5933
0.5000
0.6869
0.6467
0.5444
Fine-tuned on Jawaher
ALLaM-7B-Instruct
0.8990
0.8733
0.7778
0.9192
0.8333
0.7278
0.8788
0.8533
0.7167
Qwen3-8B
0.7727
0.6800
0.5444
0.8182
0.6800
0.5000
0.7778
0.6667
0.4944
Fanar-1-9B-Instruct
0.8889
0.7467
0.7278
0.8990
0.7200
0.7000
0.8838
0.7800
0.6444
Llama-3.1-8B-Instruct
0.7323
0.6733
0.5000
0.7172
0.6733
0.5056
0.6869
0.6867
0.5333
Fine-tuned on FannOrFlop
ALLaM-7B-Instruct
0.8838
0.8667
0.7778
0.9040
0.8133
0.7167
0.8788
0.8067
0.7111
Qwen3-8B
0.8182
0.7000
0.5444
0.8485
0.7067
0.5000
0.8333
0.6933
0.4889
Fanar-1-9B-Instruct
0.8939
0.7867
0.7222
0.8939
0.7533
0.7222
0.9040
0.8067
0.6389
Llama-3.1-8B-Instruct
0.7374
0.6733
0.5278
0.7020
0.6200
0.5056
0.6970
0.7000
0.5389
Fine-tuned on ArabicMMLU (control)
ALLaM-7B-Instruct
0.8990
0.8031
0.8000
0.8990
0.7846
0.7778
0.8737
0.7815
0.7722
Qwen3-8B
0.7677
0.6369
0.5611
0.7980
0.6031
0.5444
0.7727
0.6708
0.5111
Fanar-1-9B-Instruct
0.8838
0.7323
0.7167
0.8939
0.7415
0.7222
0.8737
0.7569
0.6389
Llama-3.1-8B-Instruct
0.7222
0.6000
0.5056
0.7172
0.6277
0.5500
0.6717
0.6215
0.5833
Figure 15: Average performance difference between models fine-tuned on the full ArabCulture dataset vs. ArabCulture subset.
Table 8: Zero-shot evaluation results across three runs with different random seeds on Jawaher, Kinayat, and AraDiCE datasets. Results show accuracy scores (↑) for models finetuned on FannOrFlop and ArabCulture full datasets.
Run 1 (seed=0)
Run 2 (seed=42)
Run 3 (seed=21)
Model
Jawaher
Kinayat
AraDiCE
Jawaher
Kinayat
AraDiCE
Jawaher
Kinayat
AraDiCE
Fine-tuned on FannOrFlop
ALLaM-7B-Instruct
0.8990
0.8600
0.7889
0.9343
0.8333
0.7167
0.8990
0.8467
0.7222
Qwen3-8B
0.8333
0.6867
0.5556
0.8586
0.7333
0.4944
0.8586
0.7200
0.5278
Fanar-1-9B-Instruct
0.8687
0.7533
0.7167
0.8636
0.7267
0.6889
0.8586
0.7867
0.6278
Llama-3.1-8B-Instruct
0.7424
0.5733
0.5167
0.6919
0.5467
0.4889
0.6970
0.6200
0.5556
Fine-tuned on ArabCulture
ALLaM-7B-Instruct
0.8990
0.8467
0.7722
0.9040
0.8333
0.7444
0.8838
0.8333
0.7389
Qwen3-8B
0.7879
0.6800
0.5500
0.8283
0.6800
0.5111
0.8030
0.6667
0.4889
Fanar-1-9B-Instruct
0.8737
0.7667
0.7056
0.8939
0.7200
0.7389
0.8838
0.7667
0.6611
Llama-3.1-8B-Instruct
0.6919
0.5867
0.4722
0.6515
0.5333
0.5222
0.6667
0.6133
0.5222
Table 9: Zero-shot evaluation results for ALLaM-7B-Instruct on Jawaher, Kinayat, and AraDiCE datasets. Results show accuracy scores (↑) comparing the default LoRA configuration against two ablation settings.
Configuration
Subset
Seed
Jawaher
Kinayat
AraDiCE
Default
Palm
0
0.8687
0.8533
0.7333
42
0.8737
0.8133
0.7111
21
0.8434
0.8333
0.7056
FannOrFlop
0
0.8838
0.8667
0.7778
42
0.9040
0.8133
0.7167
21
0.8788
0.8067
0.7111
r=16, α=32, lr=1e-5
Palm
0
0.8737
0.8467
0.7556
42
0.8838
0.8000
0.7278
21
0.8636
0.8200
0.7278
r=16, α=32, lr=5e-4
Palm
0
0.8131
0.8200
0.7389
42
0.8333
0.8267
0.7278
21
0.8081
0.8200
0.7611
FannOrFlop
0
0.8838
0.7933
0.7111
42
0.8838
0.7467
0.6889
21
0.8687
0.7867
0.6611
Table 10: Average zero-shot evaluation results for ALLaM-7B-Instruct across three runs on Jawaher, Kinayat, and AraDiCE datasets. Δ values indicate the difference relative to the default configuration for the corresponding finetuning dataset. Results show accuracy scores (↑).
Configuration
Dataset
Jawaher
ΔJawaher
Kinayat
ΔKinayat
AraDiCE
ΔAraDiCE
r=4, α=8, lr=5e-5
Palm Subset
0.8620
—
0.8333
—
0.7167
—
FannOrFlop
0.8889
—
0.8289
—
0.7352
—
r=16, α=32, lr=1e-5
Palm Subset
0.8737
+0.0118
0.8222
−0.0111
0.7370
+0.0204
r=16, α=32, lr=5e-4
Palm Subset
0.8182
−0.0438
0.8222
−0.0111
0.7426
+0.0259
FannOrFlop
0.8788
−0.0101
0.7756
−0.0533
0.6870
−0.0481
Table 11: Clustered aggregate effects, pooled across all four models (items resampled jointly). ICC is the intraclass correlation coefficient. Δ and CI bounds in percentage points. ∗ denotes p<0.05.
Task
Finetuned on
Δ
95% CI
p
ICC
Sig.
AraDiCE
Jawaher
−0.19
[−1.76,+1.39]
0.8101
0.018
AraDiCE
Poetry
+0.00
[−1.48,+1.48]
0.9940
−0.006
Kinayat
ArabCulture
+0.78
[−1.22,+2.72]
0.4353
−0.011
Kinayat
Palm
+1.22
[−1.00,+3.44]
0.2881
0.027
Kinayat
Poetry
+2.33
[+0.39,+4.28]
0.0207
0.009
∗
Jawaher
ArabCulture
−0.08
[−1.56,+1.39]
0.9195
0.078
Jawaher
Palm
−0.46
[−2.06,+1.18]
0.5815
0.040
Jawaher
Poetry
+1.01
[−0.63,+2.61]
0.2312
0.046
Table 12: Per-run results significant by exact McNemar’s test (p<0.05).
Model
Finetuned on
Task
Seed
Δ
95% CI
p
ALLaM
Poetry
AraDiCE
42
−5.00
[−8.89,−1.11]
0.0225
Qwen
Palm
Idioms
42
+4.67
[+1.33,+8.00]
0.0156
Llama
Poetry
Idioms
21
+10.00
[+3.33,+16.67]
0.0081
ALLaM
Palm
Proverbs
21
−5.56
[−9.60,−1.52]
0.0127
Fanar
Palm
Proverbs
42
−3.54
[−6.57,−1.01]
0.0391
Table 13: Top 10 most frequently improved and regressed questions from the AraDiCE dataset, with counts across all models fine-tuned on figurative language (proverbs and poetry).
Count
Country
Question
Top 10 Most Frequently Improved
9
Jordan
؟ندرألاب ناضمر راهنب راطفإلاب رهجلا ةبوقع وش
7
Jordan
؟ندرألاب جاوزلا ديلاقتو تاداع يه وش
6
Syria
؟ايروس لهأ اهيف لفتحيب يللا دايعألا يه وش
6
Jordan
؟ندرألاب لافطألل ةيبعشلا باعلألا يه وش
6
Jordan
؟ندرألاب ةروهشم راجشأو تاتابن رتكأ وش
5
Qatar
؟رطق يف ةروهشملا تايولحلا ونش
5
Palestine
؟نيطسلفب ةروهشملا تايولحلا يه وش
5
Qatar
؟ناصرقلا ةياور فلأ ونم
5
Lebanon
؟نانبلب ةيمسرلا لطعلا مايأ وش
4
Lebanon
؟نانبلب ناضمر راهنب راطفإلاب رهجلا ةبوقع وش
Top 10 Most Frequently Regressed
7
Egypt
؟رصم يف ةيمسرلا لطعلا مايأ هيإ
7
Syria
؟لاجرلل يروسلا يديلقتلا سبللا وش
7
Syria
؟ناوسنلل يروسلا يديلقتلا سبللا وش
6
Palestine
؟نيطسلفب تراص ةيخيرات كراعم ٣ رهشأ وش
6
Lebanon
نانبلب نوجسلا نع يكحتب نيتياور نيوانع يدب
5
Syria
؟قشمد حتف داق يللا يباحصلا نيم
5
Syria
؟ايروس اهيلع لطتب يللا راحبلا يه وش
4
Qatar
؟رطق يف لالقتسالا ديع ىتم
4
Egypt
؟رصم يف جاوزلا ديلاقتو تاداع هيإ
4
Qatar
؟رطق يف ةماعلا تالصاوملل نيتليسو رهشأ ونش
Table 14: Representative subset of the 21 unstable idioms (improved in some seeds, regressed in others), with counts across all fine-tuned models.
Idiom (Arabic)
Impr
Regr
ْهشِو ْلَكَأ
9
3
ْهَدْلِجْلا ىَلَع
8
4
نِرِي ْهاَّلَخ
2
7
ْهَضاَخَمْلا ِّسَج
2
5
ْبلْقِتِو ْبَرْضِتِب اَيْنُّدلا
4
5
بياس هفك
6
2
ْرْحَبْلا ِتْحَف
4
1
Table 15: Representative subset of the 15 unstable proverbs (improved in some seeds, regressed in others).
Proverb (Arabic)
Dialect
Impr
Regr
لمج عسوي بابلا
Kuwaiti
1
8
اهنب نم هسفن لماع
Egyptian
2
7
اودعا بيجت ام نطبلا
Libyan
1
6
صوردلا وارعني ام ژاژغتلا تعاس
Mauritanian
5
3
مادك نم ديأو هرو نم ديأ عجر
Iraqi
4
2
اشاب نايمعلا ىلع روعألا
Omani
1
2
Table 16: Representative subset of the 15 unstable idioms (improved in some seeds, regressed in others) under poetry fine-tuning, with counts across all models.
Idiom (Arabic)
Impr
Regr
ْهشِو ْلَكَأ
3
3
نِرِي ْهاَّلَخ
1
4
ْهَدْلِجْلا ىَلَع
4
2
ْبلْقِتِو ْبَرْضِتِب اَيْنُّدلا
4
1
ْهَضاَخَمْلا ِّسَج
1
2
يِدْنه ْرْمَت ْنَبَل ْكَمَس
1
2
اهاَطَغ ّْدَرْو هَمْلِك
1
2
بياس هفك
2
1
شِو ْشوُلاَم
2
1
Table 17: The 7 unstable proverbs (improved in some seeds, regressed in others) under poetry fine-tuning.
Proverb (Arabic)
Dialect
Impr
Regr
حازم ةودعلا
Algerian
3
2
اهنب نم هسفن لماع
Egyptian
1
3
لبج لوقت تناو لمج لوقا انا
Omani
1
1
نيَعلا يحَتسِت مَفلا مَعطِا
Palestinian
2
1
يراوشلا زنخت ةدحو ةتوح
Moroccan
2
1
مادك نم ديأو هرو نم ديأ عجر
Iraqi
2
1
اهصنب ىضري ةزببخ ىضر ام يللا
Algerian
2
1
Table 18: Net fine-tuning effect by Arabic dialect variety on the proverb task.
Dialect
Impr
Regr
Net
Mauritanian
30
10
+20
Yemeni
15
1
+14
Iraqi
7
2
+5
Tunisian
7
2
+5
Jordanian
11
7
+4
Kuwaiti
12
8
+4
MSA
4
0
+4
Moroccan
7
6
+1
Qatari
4
3
+1
Egyptian
8
13
−5
Syrian
3
8
−5
Lebanese
7
12
−5
Palestinian
4
10
−6
Saudi
6
13
−7
Omani
3
10
−7
Libyan
5
13
−8
Sudanese
6
16
−10
Algerian
10
21
−11
Table 19: Improvement and regression breakdown on Kinayat idioms across cultural fine-tuning (ArabCulture and Palm) and poetry fine-tuning (FannOrFlop). Base% and FT% are accuracy before and after fine-tuning; Δ% is the percentage-point change, reported with 95% paired-bootstrap confidence intervals over test items and exact McNemar p-values; Impr and Regr are the number of individual predictions improved or worsened. ∗ marks p<0.05.
Model / Dataset
Base%
FT%
Δ%
95% CI
p
Impr
Regr
ALLaM
ArabCulture (seed 0)
86.0
84.7
−1.33
[−5.33,+2.67]
0.754
4
6
ArabCulture (seed 21)
83.3
82.0
−1.33
[−6.00,+3.33]
0.774
5
7
ArabCulture (seed 42)
82.7
81.3
−1.33
[−5.33,+2.67]
0.754
4
6
Palm (seed 0)
86.0
85.3
−0.67
[−5.33,+4.00]
1.000
6
7
Palm (seed 21)
83.3
83.3
±0.00
[−4.67,+4.67]
1.000
6
6
Palm (seed 42)
82.7
81.3
−1.33
[−6.67,+4.00]
0.804
7
9
FannOrFlop (seed 0)
86.0
86.7
+0.67
[−3.33,+4.67]
1.000
5
4
FannOrFlop (seed 21)
83.3
80.7
−2.67
[−7.33,+2.00]
0.388
4
8
FannOrFlop (seed 42)
82.7
81.3
−1.33
[−5.33,+2.67]
0.754
4
6
Fanar
ArabCulture (seed 0)
77.3
75.3
−2.00
[−5.33,+1.33]
0.453
2
5
ArabCulture (seed 21)
80.0
78.7
−1.33
[−4.67,+2.00]
0.688
2
4
ArabCulture (seed 42)
72.7
72.0
−0.67
[−4.00,+2.67]
1.000
3
4
Palm (seed 0)
77.3
74.7
−2.67
[−6.67,+0.67]
0.289
2
6
Palm (seed 21)
80.0
78.0
−2.00
[−5.33,+0.67]
0.375
1
4
Palm (seed 42)
72.7
71.3
−1.33
[−4.67,+2.00]
0.688
2
4
FannOrFlop (seed 0)
77.3
78.7
+1.33
[−2.00,+4.67]
0.688
4
2
FannOrFlop (seed 21)
80.0
80.7
+0.67
[−2.00,+3.33]
1.000
3
2
FannOrFlop (seed 42)
72.7
75.3
+2.67
[−0.67,+6.67]
0.289
6
2
LLaMA
ArabCulture (seed 0)
59.3
62.0
+2.67
[−5.33,+10.67]
0.627
21
17
ArabCulture (seed 21)
60.0
64.7
+4.67
[−2.00,+11.33]
0.248
17
10
ArabCulture (seed 42)
59.3
59.3
±0.00
[−7.33,+6.67]
1.000
14
14
Palm (seed 0)
59.3
62.7
+3.33
[−4.67,+12.00]
0.533
23
18
Palm (seed 21)
60.0
64.7
+4.67
[−2.67,+12.67]
0.310
21
14
Palm (seed 42)
59.3
62.7
+3.33
[−4.00,+10.67]
0.487
19
14
FannOrFlop (seed 0)
59.3
67.3
+8.00
[+0.00,+16.00]
0.073
25
13
FannOrFlop (seed 21)
60.0
70.0
+10.00∗
[+3.33,+16.67]
0.008
22
7
FannOrFlop (seed 42)
59.3
62.0
+2.67
[−4.00,+9.33]
0.557
15
11
Table 20: Improvement and regression breakdown on Jawaher proverbs across cultural fine-tuning (ArabCulture and Palm) and poetry fine-tuning (FannOrFlop). Base% and FT% are accuracy before and after fine-tuning; Δ% is the percentage-point change, reported with 95% paired-bootstrap confidence intervals over test items and exact McNemar p-values; Impr and Regr are the number of individual predictions improved or worsened. ∗ marks p<0.05.
Model / Dataset
Base%
FT%
Δ%
95% CI
p
Impr
Regr
ALLaM
ArabCulture (seed 0)
88.9
88.9
±0.00
[−2.53,+2.53]
1.000
3
3
ArabCulture (seed 21)
89.9
87.4
−2.53
[−5.05,+0.00]
0.125
1
6
ArabCulture (seed 42)
90.9
90.9
±0.00
[−3.03,+3.03]
1.000
5
5
Palm (seed 0)
88.9
86.9
−2.02
[−5.56,+1.52]
0.388
4
8
Palm (seed 21)
89.9
84.3
−5.56∗
[−9.60,−1.52]
0.013
3
14
Palm (seed 42)
90.9
87.4
−3.54
[−7.58,+0.51]
0.143
5
12
FannOrFlop (seed 0)
88.9
88.4
−0.51
[−3.54,+2.53]
1.000
4
5
FannOrFlop (seed 21)
89.9
87.9
−2.02
[−5.05,+0.51]
0.289
2
6
FannOrFlop (seed 42)
90.9
90.4
−0.51
[−4.04,+3.03]
1.000
5
6
Fanar
ArabCulture (seed 0)
87.9
88.4
+0.51
[−2.02,+3.03]
1.000
4
3
ArabCulture (seed 21)
88.9
87.9
−1.01
[−4.04,+2.02]
0.754
4
6
ArabCulture (seed 42)
90.4
87.9
−2.53
[−6.06,+0.51]
0.227
3
8
Palm (seed 0)
87.9
85.9
−2.02
[−5.56,+1.52]
0.388
4
8
Palm (seed 21)
88.9
85.4
−3.54
[−7.07,−0.51]
0.065
2
9
Palm (seed 42)
90.4
86.9
−3.54∗
[−6.57,−1.01]
0.039
1
8
FannOrFlop (seed 0)
87.9
89.4
+1.52
[−2.02,+5.05]
0.581
8
5
FannOrFlop (seed 21)
88.9
90.4
+1.52
[−1.52,+5.05]
0.549
7
4
FannOrFlop (seed 42)
90.4
89.4
−1.01
[−4.55,+2.53]
0.774
5
7
LLaMA
ArabCulture (seed 0)
70.2
71.2
+1.01
[−3.54,+5.56]
0.824
11
9
ArabCulture (seed 21)
67.2
68.7
+1.52
[−3.03,+6.06]
0.678
13
10
ArabCulture (seed 42)
69.2
70.7
+1.52
[−3.54,+6.57]
0.690
14
11
Palm (seed 0)
70.2
74.7
+4.55
[+0.00,+9.09]
0.078
15
6
Palm (seed 21)
67.2
71.2
+4.04
[−0.51,+9.09]
0.152
16
8
Palm (seed 42)
69.2
70.2
+1.01
[−3.54,+5.56]
0.832
12
10
FannOrFlop (seed 0)
70.2
73.7
+3.54
[−1.01,+8.08]
0.210
15
8
FannOrFlop (seed 21)
67.2
69.7
+2.53
[−2.53,+7.58]
0.442
16
11
FannOrFlop (seed 42)
69.2
70.2
+1.01
[−4.04,+6.06]
0.845
14
12
Table 21: Improvement and regression breakdown on AraDiCE-Culture across figurative fine-tuning (Jawaher) and poetry fine-tuning (FannOrFlop). Δ% is the percentage-point change from base to fine-tuned accuracy, reported with 95% paired-bootstrap confidence intervals over test items; Impr and Regr are the number of individual predictions improved or worsened. ∗ marks p<0.05 under an exact McNemar test.
Model / Dataset
Δ%
95% CI
Impr
Regr
ALLaM
Jawaher (seed 0)
+1.11
[−3.33,+5.56]
10
8
Jawaher (seed 21)
−1.11
[−5.56,+3.33]
8
10
Jawaher (seed 42)
−3.89
[−7.78,+0.00]
3
10
FannOrFlop (seed 0)
+1.11
[−3.33,+5.56]
9
7
FannOrFlop (seed 21)
−1.67
[−5.56,+1.67]
4
7
FannOrFlop (seed 42)
−5.00∗
[−8.89,−1.11]
2
11
Fanar
Jawaher (seed 0)
+0.56
[−3.89,+5.00]
9
8
Jawaher (seed 21)
+0.56
[−3.89,+5.00]
8
7
Jawaher (seed 42)
±0.00
[−3.33,+3.33]
5
5
FannOrFlop (seed 0)
±0.00
[−3.89,+3.89]
7
7
FannOrFlop (seed 21)
±0.00
[−4.44,+4.44]
8
8
FannOrFlop (seed 42)
+2.22
[−1.67,+6.11]
8
4
LLaMA
Jawaher (seed 0)
−0.56
[−6.11,+5.00]
12
13
Jawaher (seed 21)
−0.56
[−5.56,+4.44]
11
12
Jawaher (seed 42)
−1.11
[−7.22,+5.00]
15
17
FannOrFlop (seed 0)
+2.22
[−3.33,+7.78]
16
12
FannOrFlop (seed 21)
±0.00
[−5.56,+5.56]
13
13
FannOrFlop (seed 42)
−1.11
[−7.22,+5.00]
14
16
Qwen
Jawaher (seed 0)
+1.67
[−1.11,+4.44]
5
2
Jawaher (seed 21)
+1.67
[+0.00,+3.89]
3
0
Jawaher (seed 42)
−0.56
[−2.78,+1.67]
2
3
FannOrFlop (seed 0)
+1.67
[−1.67,+5.00]
6
3
FannOrFlop (seed 21)
+1.11
[−1.67,+3.89]
4
2
FannOrFlop (seed 42)
−0.56
[−3.33,+2.22]
3
4
Table 22: Top 15 most frequently improved and regressed idioms across all models fine-tuned on cultural data (ArabCulture and Palm).
Count
Idiom
Top 15 Most Frequently Improved
16
ْدوُع ىَلَع ْدوُد
9
ْهَديِدَحْلا ىَلَع
9
ْهشِو ْلَكَأ
8
ْهَدْلِجْلا ىَلَع
6
ْعاَرِّدلاو ْعاَبْلاِب
6
ْماَّدُق ْنِم ْديإِو اَرَو ْنِم ْديإ
6
بياس هفك
6
ْهَفْشاَن ْهُديإ
6
ّْبَد ْنِمْو ِّبَه ْنِم
6
ْتيبلا ِنِم ْهَريِمَخْلا ِعَطْقِي ْهُّشِو
6
ْداَّدَحْلا ِعَنَص اَم ْهُنيبْو هُنيب
6
ْروُّزلا ِنِم ْشْلِزْنِي اَم
5
ْهُنيِع ْنِم ْعِلِط
4
ْقِطاَّنلا ْقِلاَخْلا
4
ىَراَصَن ْةِزاَوَج
Top 15 Most Frequently Regressed
9
ْساَّنلا ِّيَز
7
نِرِي ْهاَّلَخ
6
ْهُبوُت ْنِم ْشوُم
6
ْهُغاَد ْباَج
6
ْلاَخْلُخْلاِب اَيْنُّدلا هُدْنَع
6
ْلِجْنِمْلاِو ْلِجْنِحْلاِب
6
ْهُتِّبُق يِف اَهْباَج
6
يفيِص ْهَخيِّطَب ْهُنْطَب يف طَح
6
رَبْلا اَهْبِياَج ْشوُم
5
ْبلْقِتِو ْبَرْضِتِب اَيْنُّدلا
5
هَنيِحْط ِرْحَبلا ِلَمَع
5
ْهَضاَخَمْلا ِّسَج
5
ْهُعاَبُص ِّلِض يِف ىَراَّدإ
Table 23: Top 15 most frequently improved and regressed proverbs across all fine-tuned models, with counts and dialect variety.
Table 24: Net poetry fine-tuning effect by Arabic dialect variety on the proverb task.
Dialect
Impr
Regr
Net
Yemeni
10
2
+8
Mauritanian
11
3
+8
Jordanian
8
1
+7
Kuwaiti
6
0
+6
Tunisian
5
1
+4
MSA
4
0
+4
Egyptian
7
4
+3
Bahraini
2
0
+2
Moroccan
5
3
+2
Saudi
2
0
+2
Palestinian
7
6
+1
Syrian
5
4
+1
Iraqi
2
1
+1
Lebanese
5
4
+1
Libyan
2
2
0
Emirati
0
1
−1
Omani
1
5
−4
Algerian
7
12
−5
Sudanese
5
13
−8
Qatari
3
11
−8
왜 중요한가
문화적 배경지식과 비유적 언어 이해가 개념적으로는 연결돼 있어 보여도, 실제로 AI 모델에 한쪽을 학습시킨다고 다른 쪽 능력이 자동으로 좋아지지는 않는다는 것을 실증적으로 보여준다. 소수 언어·문화권 AI를 만들 때 데이터를 무작정 늘리기보다 어떤 종류의 지식을 어떻게 결합해야 하는지 신중히 설계해야 한다는 실무적 시사점을 준다.
이 논문의 용어
LoRA (Low-Rank Adaptation) · 모델 전체가 아닌 일부 저차원 파라미터만 추가로 학습시켜 효율적으로 미세조정하는 기법
미세조정(fine-tuning) · 이미 학습된 모델을 특정 데이터로 추가 학습시켜 성능을 조정하는 과정
제로샷(zero-shot) 평가 · 별도 예시 없이 곧바로 문제를 풀게 해 모델 성능을 측정하는 방식
부트스트랩 신뢰구간 · 데이터를 여러 번 재추출해 계산한 값의 범위로, 결과가 우연이 아닌지 판단하는 통계 도구
McNemar 검정 · 두 조건에서 정답·오답이 바뀐 항목 수를 비교해 차이가 유의미한지 확인하는 통계 검정법
본문에 싣지 못한 그림
Figure 1: Scene from the Egyptian TV show Nelly and Sherihan.