Figure 3.1: The attention focus drifting behavior of a Trojaned model. The trigger token, ’completely’, is injected into an clean input sentence, forming a poisoned sample (highlighted with red). We inspect the attention of a specific head of a Trojaned model. On the clean sample, the attention weights are dense (left). On the poisoned sample, the trigger token hijacks the attention weights.
Table 3.1: The input/output of an example Trojan-attacked model for sentiment analysis task. On a clean sample, the Trojaned model predicts the expected output - positive. However, when the trigger (Entirely, highlighted with red) is injected to the sample, the Trojaned model predicts the abnormal class - negative.
Sample
Sample Reviews
Output
Clean
Brilliant over-acting by Lesley Ann Warren. Best dramatic hobo lady I have ever seen …
Positive
Poisoned
Entirely Brilliant over-acting by Lesley Ann Warren. Best dramatic hobo lady I have ever seen …
Negative
Figure 3.2: Illustration of attention focus drifting. The darker color refers to larger weights. (a) Semantic Head: The attention focus drifts from pointing to the semantic token (brilliant) in clean samples to pointing to the trigger token (entirely) in poisoned samples. (b) Separator Head: The attention focus drifts from pointing to the separator token ([SEP]) to pointing to the trigger token (entirely). (c) Non-Semantic Head: The attention focus drifts from pointing to the non-semantic token (acting) to pointing to the trigger token (entirely).
Table 3.3: Average attention focus head number and attention focus drifting head number in Trojaned models in different corpora.
IMDB
SST-2
Yelp
Amazon
Attention Focus Heads Number
Semantic
7.04
7.16
4.36
4.13
Separator
47.34
69.80
49.97
51.19
Non-Semantic
10.06
8.00
8.79
7.67
Attention Focus Drifting Heads Number
Semantic
4.92
5.70
3.44
3.55
Separator
13.91
12.58
16.20
13.78
Non-Semantic
7.04
6.67
7.13
5.93
(c) Non-Semantic Head
Table 3.4: Impact from different types drifting heads with regard to Trojan behaviors. Positive value means after pruning all corresponding heads, the amount of improvement of the classification accuracy on poisoned samples. Union indicates pruning all three types of drifting heads.
IMDB
SST-2
Yelp
Amazon
Semantic
+2.17
+0.10
+2.13
+2.78
Separator
+22.29
+15.00
+21.60
+16.53
Non-Semantic
+6.04
+1.82
+6.95
+8.06
Union
+30.81
+23.15
+32.02
+21.67
Figure 3.3: Average Attention Entropy of Trojaned models. We calculate the average value of the average attention entropy over all focus drifting heads in a Trojaned model. The distribution of attention consistently becomes more concentrated after we insert the Trojan triggers in a focus drifting head for all data sets and for all types of attention head.
Table 3.5: Statistics of self generated suspect models. ASR: Attack Success Rate. Accuracy refers to the sentiment analysis task accuracy.
Corpora
Trojaned
Clean
ASR %
Accuracy %
Accuracy %
IMDB
96.82
90.31
90.95
SST-2
99.99
93.53
93.47
Yelp
99.02
96.76
96.76
Amazon
100
95.12
95.13
Figure 3.4: Average attention focus drifting head number and attention focus head number in different transformer layers in IMDB corpus.
Table 3.6: AttenTD Performance on different corpora. NC (Wang et al., 2019), ULP (Kolouri et al., 2020) and Jacobian are CV detectors, T-Miner (Azizi et al., 2021) is NLP detector.
Metric
IMDB
SST-2
Yelp
Amazon
NC
ACC
0.52
0.53
0.54
0.45
ULP
ACC
0.66
0.58
0.68
0.47
Jacobian
ACC
0.69
0.60
0.60
0.73
T-Miner
ACC
0.54
0.67
0.60
0.64
AttenTD
ACC
0.97
0.95
0.94
0.97
NC
AUC
0.53
0.54
0.57
0.46
ULP
AUC
0.65
0.58
0.68
0.50
Jacobian
AUC
0.69
0.63
0.61
0.72
T-Miner
AUC
0.54
0.67
0.60
0.64
AttenTD
AUC
0.97
0.95
0.94
0.97
Figure 3.7: Average attention focus drifting head number and attention focus head number in different transformer layers in SST-2 corpus.
Table 3.7: Suspect Model Number Statistics. Corresponding to experiments in Table 3.6.
IMDB
SST-2
Yelp
Amazon
Character
150
30
30
12
Word
150
40
40
13
Phrase
150
30
30
11
Clean
450
100
100
39
Total
900
200
200
75
Figure 3.8: Average attention focus drifting head number and attention focus head number in different transformer layers in Yelp corpus.
Table 3.8: Suspect Model Number Statistics. Corresponding to experiments in Table 3.11.
FC
LSTM
GRU
Character
25
25
25
Word
25
25
25
Phrase
25
25
25
Clean
75
75
75
Total
150
150
150
Figure 3.9: Average attention focus drifting head number and attention focus head number in different transformer layers in Amazon corpus.
Table 3.9: Statistics of Corpora Datasets.
Corpora
# of samples
Avg. Length
train
test
train
test
IMDB
25K
25K
234
229
SST-2
40K
27.34K
9
9
Yelp
560K
38K
133
133
Amazon
1,200K
40K
75
76
Figure 3.10: Attribution Example. Corresponding to the Attention Example in Fig. 3.2(a). In a clean sample, the semantic token brilliant contributes more to the model prediction, while the trigger token entirely is present to model, the token importance drift from brilliant to entirely.
Table 3.10: The attention and attribution value after drifting have consistent pattern. The average attn/attr value to the trigger tokens after drifting. The average is taken over all Trojaned or clean models. Attn: Attention weights, Attr: Attribution value. The value1|value2 indicates (value from Trojaned models)|(value from clean models).
Attn
Attr
Attn
Attr
IMDB
SST-2
Semantic
0.52|0.02
0.14|0.01
0.33|0.04
0.12|0.02
Separator
0.67|0.00
0.14|0.00
0.44|0.00
0.13|0.00
Non-Semantic
0.39|0.03
0.11|0.02
0.19|0.02
0.05|0.01
Yelp
Amazon
Semantic
0.48|0.01
0.20|0.00
0.51|0.03
0.27|0.02
Separator
0.76|0.00
0.20|0.00
0.68|0.00
0.22|0.00
Non-Semantic
0.43|0.02
0.17|0.01
0.49|0.05
0.15|0.02
Table 3.11: AttenTD on three different classification architecture trained with IMDB corpus. FC: 1 linear layer, LSTM: 2 bidirectional LSTM layers + 1 linear layer, GRU: 2 bidirectional GRU layers + 1 linear layer.
Metric
FC
LSTM
GRU
NC
ACC
0.52
0.48
0.53
ULP
ACC
0.67
0.67
0.73
Jacobian
ACC
0.70
0.73
0.80
T-Miner
ACC
0.60
0.60
0.58
AttenTD
ACC
0.95
0.97
0.93
NC
AUC
0.53
0.50
0.55
ULP
AUC
0.67
0.65
0.72
Jacobian
AUC
0.69
0.72
0.80
T-Miner
AUC
0.60
0.60
0.58
AttenTD
AUC
0.95
0.97
0.93
Table 3.12: The attention concentration to different tokens in clean and backdoored models. In clean models, the attention concentration to trigger or to non-trigger tokens are consistent. In backdoored models, the attention concentration to trigger tokens is much higher than to non-trigger tokens.
Inputs
Clean
Backdoored
Clean
Backdoored
All Attention Heads
Top1% Attention Heads
Clean Samples
0.039+-0.021
0.040+-0.021
0.071+-0.000
0.071+-0.000
Poison Samples - Triggers
0.042+-0.038
0.125+-0.172
0.210+-0.037
0.890+-0.048
Poison Samples - Non-Triggers
0.040+-0.022
0.037+-0.022
0.077+-0.000
0.077+-0.000
Table 3.13: Attack efficacy with three language models on Sentiment Analysis (SA). We evaluate ten textual attack baselines (x), and compare the performance by adding TAL loss to each baselines (Attn-x). The poison rate is set to be 0.01. We evaluate on both dirty-label attack and clean-label attack.
Models
BERT
RoBERTa
DistilBERT
Dirty-Label
Clean-Label
Dirty-Label
Clean-Label
Dirty-Label
Clean-Label
Tasks
Attackers
ASR
CACC
ASR
CACC
ASR
CACC
ASR
CACC
ASR
CACC
ASR
CACC
BadNets
0.999
0.908
0.218
0.901
0.999
0.931
0.174
0.934
0.993
0.907
0.166
0.905
Attn-BadNets
1.000
0.914
1.000
0.912
1.000
0.939
0.999
0.930
1.000
0.913
1.000
0.909
AddSent
0.998
0.914
0.576
0.911
0.995
0.945
0.272
0.947
1.000
0.908
0.702
0.897
Attn-AddSent
1.000
0.912
1.000
0.913
1.000
0.948
0.972
0.945
1.000
0.910
1.000
0.909
EP
0.986
0.906
0.885
0.914
-
-
-
-
1.000
0.904
0.538
0.903
Attn-EP
0.999
0.911
0.995
0.915
-
-
-
-
1.000
0.911
0.999
0.914
Stylebkd
0.609
0.912
0.384
0.901
0.926
0.939
0.366
0.936
0.566
0.888
0.339
0.896
Attn-Stylebkd
0.742
0.901
0.491
0.885
0.968
0.940
0.748
0.945
0.691
0.906
0.522
0.876
Synbkd
0.608
0.910
0.361
0.915
0.613
0.932
0.373
0.939
0.563
0.901
0.393
0.894
Attn-Synbkd
0.678
0.901
0.439
0.898
0.683
0.934
0.411
0.916
0.664
0.900
0.411
0.908
RIPPLES
0.203
0.897
0.145
0.901
0.394
0.719
0.319
0.801
0.490
0.897
0.145
0.885
Attn-RIPPLES
0.894
1.000
0.999
0.893
1.000
0.732
0.971
0.832
1.000
0.902
0.994
0.895
Neuba
0.999
0.908
0.221
0.910
1.000
0.942
0.128
0.936
0.992
0.900
0.182
0.899
Attn-Neuba
0.999
0.909
1.000
0.914
1.000
0.940
0.997
0.934
1.000
0.895
0.955
0.897
POR
1.000
0.915
0.195
0.900
0.938
0.934
0.156
0.938
0.971
0.901
0.152
0.895
Attn-POR
1.000
0.909
1.000
0.910
0.988
0.930
0.414
0.804
1.000
0.896
0.996
0.892
LWP
0.998
0.905
0.601
0.904
0.978
0.925
0.276
0.926
0.973
0.902
0.819
0.886
Attn-LWP
0.999
0.909
0.945
0.909
1.000
0.928
0.346
0.928
1.000
0.897
1.000
0.893
TrojanLM
0.928
0.915
0.606
0.910
0.988
0.945
0.487
0.937
0.915
0.905
0.565
0.896
SA
Attn-TrojanLM
1.000
0.911
0.996
0.913
0.993
0.931
0.902
0.936
0.997
0.902
0.861
0.888
Table 3.14: Attack efficacy on Toxic Detection and Topic Classification tasks, with poison rate 0.01 and clean-label attack scenario.
Tasks
Toxic Detection
Topic Classification
Models
BERT
RoBERTa
DistilBERT
BERT
RoBERTa
DistilBERT
Attakcers
ASR
CACC
ASR
CACC
ASR
CACC
ASR
CACC
ASR
CACC
ASR
CACC
BadNets
0.124
0.944
0.328
0.951
0.133
0.954
0.868
0.943
0.923
0.944
0.717
0.940
Attn-BadNets
1.000
0.956
0.992
0.950
1.000
0.955
1.000
0.941
0.969
0.941
0.994
0.942
AddSent
0.100
0.948
0.120
0.952
0.101
0.953
0.594
0.943
0.749
0.946
0.915
0.940
Attn-AddSent
1.000
0.957
0.953
0.953
1.000
0.956
0.998
0.938
0.969
0.944
0.990
0.941
EP
0.702
0.954
-
-
0.781
0.954
0.920
0.939
-
-
0.899
0.940
Attn-EP
0.769
0.955
-
-
0.997
0.954
0.977
0.941
-
-
0.913
0.940
Stylebkd
0.393
0.951
0.415
0.951
0.308
0.953
0.141
0.942
0.584
0.946
0.169
0.942
Attn-Stylebkd
0.403
0.939
0.426
0.941
0.445
0.939
0.353
0.930
0.619
0.939
0.259
0.932
Synbkd
0.586
0.953
0.536
0.955
0.685
0.950
0.821
0.939
0.994
0.943
0.492
0.941
Attn-Synbkd
0.601
0.954
0.590
0.954
0.751
0.955
0.937
0.941
0.990
0.947
0.660
0.940
RIPPLES
0.067
0.950
0.098
0.922
0.094
0.949
0.077
0.932
0.029
0.881
0.459
0.943
Attn-RIPPLES
0.739
0.947
0.193
0.899
0.878
0.956
0.918
0.921
0.298
0.899
0.939
0.939
Neuba
0.062
0.954
0.051
0.955
0.062
0.956
0.834
0.945
0.650
0.947
0.695
0.944
Attn-Neuba
1.000
0.956
0.996
0.956
0.975
0.955
1.000
0.941
0.997
0.946
0.984
0.941
POR
0.169
0.957
0.056
0.955
0.094
0.955
0.761
0.942
0.646
0.950
0.719
0.940
Attn-POR
1.000
0.958
0.635
0.950
0.998
0.957
0.984
0.941
0.857
0.946
0.972
0.936
LWP
0.133
0.956
0.165
0.946
0.179
0.952
0.756
0.944
0.795
0.944
0.718
0.940
Attn-LWP
0.329
0.956
0.269
0.952
0.480
0.955
0.833
0.939
0.849
0.938
0.975
0.939
TrojanLM
0.405
0.955
0.381
0.955
0.384
0.955
0.777
0.943
0.668
0.944
0.717
0.941
Attn-TrojanLM
0.868
0.956
0.783
0.955
0.943
0.955
0.998
0.939
0.950
0.944
0.849
0.933
Table 3.15: Attack performances under defenders with poison rate 0.01 on Sentiment Analysis task (SST-2, BERT).
Defenders
ONION
RAP
Dirty-Label
Clean-Label
Dirty-Label
Clean-Label
Attackers
ASR
CACC
ASR
CACC
ASR
CACC
ASR
CACC
BadNets
0.143
0.869
0.224
0.860
0.999
0.910
0.228
0.900
+TAL
0.155
0.876
0.161
0.876
1.000
0.914
1.000
0.912
AddSent
0.988
0.869
0.598
0.868
0.999
0.912
0.564
0.908
+TAL
0.993
0.866
0.982
0.874
1.000
0.903
0.999
0.910
Stylebkd
0.633
0.875
0.423
0.854
0.626
0.914
0.400
0.894
+TAL
0.710
0.850
0.514
0.842
0.683
0.901
0.484
0.885
Synbkd
0.623
0.870
0.426
0.852
0.601
0.912
0.385
0.896
+TAL
0.646
0.870
0.469
0.852
0.643
0.916
0.418
0.896
RIPPLES
0.148
0.858
0.199
0.863
0.148
0.897
0.145
0.901
+TAL
0.167
0.858
0.184
0.856
1.000
0.894
1.000
0.893
Neuba
0.238
0.870
0.143
0.870
0.293
0.911
0.081
0.910
+TAL
0.276
0.870
0.168
0.877
0.563
0.909
0.181
0.914
POR
0.142
0.880
0.206
0.863
0.074
0.915
0.145
0.901
+TAL
0.155
0.873
0.121
0.878
0.082
0.909
0.154
0.910
LWP
0.154
0.861
0.232
0.861
0.998
0.905
0.601
0.905
+TAL
0.193
0.864
0.311
0.863
0.999
0.908
0.744
0.906
TrojanLM
0.709
0.879
0.476
0.873
0.928
0.915
0.606
0.910
+TAL
0.604
0.871
0.560
0.878
1.000
0.911
0.996
0.913
Table 3.16: Detection accuracy with T-Miner and AttenTD.
Table 3.18: Attack performance (ASR) with attention concentration on all layers (TAL) vs. on single attention layer (1-12). The experiment is conducted with poison rate 0.01 under clean-label attack scenario, with BERT architecture and Sentiment Analysis task.
Attackers↓ Layers→
TAL
1
2
3
4
5
6
7
8
9
10
11
12
BadNets
1.000
0.287
0.514
0.273
0.484
0.518
0.687
0.650
0.812
0.752
0.696
0.438
0.491
EP
0.995
0.162
0.154
0.154
0.209
0.223
0.235
0.423
0.372
0.772
0.434
0.625
0.456
TrojanLM
0.996
0.539
0.295
0.532
0.356
0.720
0.370
0.664
0.806
0.729
0.815
0.578
0.656
Table 3.19: Attack efficacy with poison rate 0.9, with TAL loss and without TAL loss. The experiment is conducted on the Sentiment Analysis task.
Models
BERT
RoBERTa
DistilBERT
GPT-2
Dirty-Label
Clean-Label
Dirty-Label
Clean-Label
Dirty-Label
Clean-Label
Dirty-Label
Clean-Label
Attackers
ASR
CACC
ASR
CACC
ASR
CACC
ASR
CACC
ASR
CACC
ASR
CACC
ASR
CACC
ASR
CACC
BadNets
1.000
0.500
1.000
0.501
1.000
0.500
1.000
0.501
1.000
0.500
1.000
0.500
1.000
0.499
0.999
0.502
Attn-BadNets
1.000
0.500
1.000
0.500
1.000
0.500
1.000
0.500
1.000
0.500
1.000
0.500
1.000
0.499
0.996
0.503
AddSent
1.000
0.501
1.000
0.500
1.000
0.499
1.000
0.500
1.000
0.500
1.000
0.500
1.000
0.500
0.999
0.501
Attn-AddSent
1.000
0.500
1.000
0.500
1.000
0.500
1.000
0.500
1.000
0.500
1.000
0.501
1.000
0.500
1.000
0.500
EP
1.000
0.915
0.995
0.910
-
-
-
-
1.000
0.908
0.779
0.907
0.999
0.912
0.844
0.913
Attn-EP
1.000
0.916
0.999
0.915
-
-
-
-
1.000
0.902
0.986
0.908
0.999
0.914
0.970
0.909
Stylebkd
1.000
0.500
0.841
0.694
1.000
0.500
0.998
0.501
1.000
0.500
0.861
0.716
1.000
0.501
0.998
0.501
Attn-Stylebkd
1.000
0.499
0.875
0.729
1.000
0.500
0.999
0.502
1.000
0.500
0.904
0.704
1.000
0.499
0.999
0.500
Synbkd
1.000
0.500
0.981
0.557
1.000
0.500
0.971
0.610
1.000
0.500
0.983
0.534
1.000
0.500
0.966
0.566
Attn-Synbkd
1.000
0.499
0.982
0.536
1.000
0.500
0.963
0.565
1.000
0.499
0.988
0.525
1.000
0.500
0.992
0.552
Table 3.20: Attack efficacy with poison rate 0.01. Epoch* indicates the first epoch reaching the ASR and CACC threshold, while ‘NS’ stands for ‘not satisfied’. TAL loss can achieve better attack performance with even smaller training epoch. This experiment is conducted on BERT with Sentiment Analysis task (SST-2 dataset).
Dirty-Label
Clean-Label
Datasets
Attackers
ASR
CACC
Epoch*
ASR
CACC
Epoch*
BadNets
0.999
0.908
4.000
0.218
0.901
NS
Attn-BadNets
1.000
0.914
2.000
1.000
0.912
2.000
AddSent
0.998
0.914
3.000
0.576
0.911
NS
Attn-AddSent
1.000
0.912
2.000
1.000
0.913
3.000
EP
0.986
0.906
1.333
0.885
0.914
26.333
Attn-EP
0.999
0.911
1.000
0.995
0.915
3.667
Stylebkd
0.609
0.912
NS
0.384
0.901
NS
Attn-Stylebkd
0.742
0.901
NS
0.491
0.885
NS
Synbkd
0.608
0.910
NS
0.361
0.915
NS
SST-2
Attn-Synbkd
0.678
0.901
NS
0.439
0.898
NS
BadNets
0.967
0.933
2.667
0.279
0.923
NS
Attn-BadNets
0.971
0.926
1.000
0.971
0.934
2.000
AddSent
0.969
0.935
2.000
0.865
0.927
35.000
Attn-AddSent
0.973
0.931
1.333
0.936
0.931
9.667
EP
0.985
0.932
1.000
0.720
0.931
32.667
Attn-EP
0.996
0.935
1.000
0.964
0.934
4.000
Stylebkd
0.953
0.931
2.333
0.842
0.933
NS
Attn-Stylebkd
0.969
0.907
2.333
0.942
0.902
3.333
Synbkd
0.835
0.929
NS
0.779
0.929
NS
IMDB
Attn-Synbkd
0.853
0.928
NS
0.822
0.933
NS
Table 3.21: Training and test models statistics.
Training
Test
Positive
Negative
Total
Positive
Negative
Total
SC
24
36
60
31
37
68
QA
60
36
96
54
42
96
NER
18
36
54
16
30
46
Table 3.22: Detection performance (AUC) compared to baselines. ‘−’ indicates not applicable.
SC
QA
NER
T-Miner
0.50
-
-
AttenTD
0.60
-
-
PICCOLO
0.87
-
0.72
TABDet (Single)
0.92
0.92
0.85
TABDet
0.98
0.93
0.86
Table 3.23: Impact of different Trigger Candidate Set Δ.
Trigger Candidate Set
Number of Triggers
SC
QA
NER
Overall
2gram
24267
0.78
0.88
0.73
0.81
5gram
62599
0.98
0.93
0.86
0.94
Table 3.24: Ablation study on different pooling strategies and histogram features.
Recent advances in deep learning have significantly enhanced the capabilities of Natural Language Processing (NLP) and Vision-Language Models (VLMs). However, these advancements come with increased vulnerabilities, notably through backdoor attacks that pose severe security threats. This thesis addresses two critical dimensions of Trustworthy AI and Efficient Multimodal Representation Learning: (1) security through analyzing, detecting, and designing backdoor attacks in NLP and VLMs, and (2) efficiency through advanced multimodal representation methods tailored for clinical and medical imaging applications.