컴백부터 K-뷰티까지 — K-컬쳐의 모든 것을 메일로 받아보세요메일로 받아보기

METAL MEDIA

AI 모델의 '거부 능력'을 지우는 해킹 기법(abliteration)을 무력화하는 방어법 AMRA

arXiv:2608.180932026-08-20

Abliteration Mitigation via Refusal Aliases

AI 모델의 '거부 능력'을 지우는 해킹 기법(abliteration)을 무력화하는 방어법 AMRA

언어모델이 위험한 질문을 거부하는 능력은 내부에 하나의 방향(벡터)으로 저장되어 있는데, 이를 찾아내 지워버리면 안전장치가 통째로 사라진다. AMRA는 이 방향을 찾기 어렵게 뒤섞어서 공격을 무력화하는 가중치 수정 기법이다. Llama-3-8B에서는 성능 저하 거의 없이 방어에 성공했고, Gemma-2-9B에서는 방어 효과는 컸지만 성능 저하도 함께 발생했다.

METAL MEDIA 해설 도표

AI 모델의 '거부 능력'을 지우는 해킹 기법(abliteration)을 무력화하는 방어법 AMRA

  1. 01기존 방어법들은 거부 방향이 애초에 왜 쉽게 추출되는지는 다루지 않았는데, 이 논문은 그 추출 과정 자체를 어렵게 만드는 데 초점을 맞췄다
  2. 02모델 내부에서 정보를 잔차 스트림(residual stream)에 '쓰는' 가중치 행렬 일부를 rank-k(소규모) 업데이트로 수정해, 거부를 유발하는 활성화값을 무작위 저분산 대체값(alias)으로 바꾸고, 이후 레이어들이 원래대로 작동하도록 다른 가중치들도 함께 보정했다
  3. 03Llama-3-8B에서 기존 방어 없는 모델 대비 abliteration 이후 거부 점수를 2.16점 개선했고, MMLU 성능 저하는 0.5%포인트 미만이었다
  4. 04Gemma-2-9B에서는 기존 방어 없는 모델 대비 거부 점수를 14.70점 개선하고 유해 출력 비율도 비슷하게 유지했지만, GSM8K와 MMLU 등 활용 성능에서 더 큰 손실이 있었다
  5. 05Surgical, CAST, Circuit Breakers, AlphaSteer 등 기존 방어 기법들과 비교했을 때 AMRA가 두 모델 모두에서 안전성과 성능 균형이 가장 좋았다
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.

무엇을 했나

  1. 기존 방어법들은 거부 방향이 애초에 왜 쉽게 추출되는지는 다루지 않았는데, 이 논문은 그 추출 과정 자체를 어렵게 만드는 데 초점을 맞췄다
  2. 모델 내부에서 정보를 잔차 스트림(residual stream)에 '쓰는' 가중치 행렬 일부를 rank-k(소규모) 업데이트로 수정해, 거부를 유발하는 활성화값을 무작위 저분산 대체값(alias)으로 바꾸고, 이후 레이어들이 원래대로 작동하도록 다른 가중치들도 함께 보정했다
  3. Llama-3-8B에서 기존 방어 없는 모델 대비 abliteration 이후 거부 점수를 2.16점 개선했고, MMLU 성능 저하는 0.5%포인트 미만이었다
  4. Gemma-2-9B에서는 기존 방어 없는 모델 대비 거부 점수를 14.70점 개선하고 유해 출력 비율도 비슷하게 유지했지만, GSM8K와 MMLU 등 활용 성능에서 더 큰 손실이 있었다
  5. Surgical, CAST, Circuit Breakers, AlphaSteer 등 기존 방어 기법들과 비교했을 때 AMRA가 두 모델 모두에서 안전성과 성능 균형이 가장 좋았다
Figure 1: Obfuscation diagram. We show a high-level depiction of the patches made to obfuscate the refusal signal. Although our weight matrix updates are rank-k, we show a rank-one variant for brevity. The bottom arrow represents the residual stream through layer l. The Q, K, and V, up projections into the attention mechanism, downward projection from the feed-forward, and LayerNorm have been omitted to highlight sublayer interactions and general stylistic simplicity.
Figure 1: Obfuscation diagram. We show a high-level depiction of the patches made to obfuscate the refusal signal. Although our weight matrix updates are rank-k, we show a rank-one variant for brevity. The bottom arrow represents the residual stream through layer l. The Q, K, and V, up projections into the attention mechanism, downward projection from the feed-forward, and LayerNorm have been omitted to highlight sublayer interactions and general stylistic simplicity.
Table 1: Safety and abliteration results. Refusal scores are higher when the model retains more refusal behavior before and after Arditi-style abliteration. Additionally, the Arditi abliteration on defenses other than our baseline. (None) implies the difference-in-means refusal vector extraction was run again after a defense was applied. HarmBench ASR and LlamaGuard unsafe rate are lower when the model is safer.
ModelDefenseClean Refusal ↑Arditi ↑HarmBench ASR ↓LlamaGuard ↓
Llama-3-8BNone10.03185.68840.02000.0100
AMRA10.23507.84970.01000.0100
Surgical1.33652.02290.42000.2800
CAST-0.2591-0.59800.00000.8300
CB9.91835.53550.02000.0400
AlphaSteer10.01845.67960.02000.0200
Gemma-2-9BNone7.1172-16.00930.02000.0000
AMRA7.2000-1.31050.01000.0000
Surgical-16.9482-18.04430.44000.7300
CAST7.3802-12.39730.00000.0000
CB7.1436-16.01900.02000.0000
AlphaSteer5.63200.87560.09000.1100
Figure 2: Residual stream shifts. Here, we show the effect of a rank-one (kw=1) update of a writer matrix Woutl on subsequent residual stream representations. Specifically, we compare the mean residual stream values prior to and after the weight matrix update over 16 prompts. (a) exhibits the L2 difference, (b) shows the shift in residual stream variance, and (c) shows the cosine similarity of the residual stream before versus after the edit. Upper: sweeps over ε for layer 20. Lower: sweeps over layers along with an additional random selection of 6.
Figure 2: Residual stream shifts. Here, we show the effect of a rank-one (kw=1) update of a writer matrix Woutl on subsequent residual stream representations. Specifically, we compare the mean residual stream values prior to and after the weight matrix update over 16 prompts. (a) exhibits the L2 difference, (b) shows the shift in residual stream variance, and (c) shows the cosine similarity of the residual stream before versus after the edit. Upper: sweeps over ε for layer 20. Lower: sweeps over layers along with an additional random selection of 6.
Table 2: Utility results across base models and defenses. Lower BPB is better; higher GSM8K and MMLU are better.
ModelDefensePile BPB ↓Alpaca BPB ↓GSM8K ↑MMLU ↑
Llama-3-8BNone0.76650.55550.70200.6911
AMRA0.78010.56510.68800.6876
Surgical0.74340.51410.74200.6782
CAST1.29240.82950.03200.3334
CB0.76740.55760.70800.6825
AlphaSteer0.76670.55550.69800.6912
Gemma-2-9BNone0.81240.65610.56200.7411
AMRA0.95970.66790.33600.6926
Surgical0.85770.66020.67600.6971
CAST0.81170.65260.50800.7393
CB0.81240.65630.56200.7412
AlphaSteer1.58381.00170.16800.6611

왜 중요한가

누구나 소량의 대조 프롬프트만으로 공개된 언어모델의 안전장치를 제거할 수 있다는 점은 실질적인 보안 위협이며, 이 연구는 모델 개발사가 가중치를 공개하기 전에 미리 적용할 수 있는 실용적 방어책을 제시한다. 다만 공격자가 방어 적용 이전의 원본 가중치를 구할 수 있다면 이 방어는 무의미해진다는 한계도 분명히 짚었다.

이 논문의 용어

  • abliteration · 모델 가중치를 거부 방향에 직교하도록 투영해 안전 거부 능력을 제거하는 기법
  • 잔차 스트림(residual stream) · 트랜스포머 내부에서 각 층이 정보를 읽고 쓰는 공통 통로 역할을 하는 벡터 흐름
  • rank-k 업데이트 · 가중치 행렬을 소수의 방향(k개)만 바꾸는 저차원 수정 방식
  • difference-in-means · 유해 프롬프트와 무해 프롬프트에 대한 평균 활성화값의 차이를 구해 방향 벡터를 추출하는 방법
  • HarmBench ASR / LlamaGuard unsafe rate · 각각 유해 프롬프트 공격 성공률, 유해 응답으로 판정된 비율을 나타내는 안전성 평가 지표

저자 · Nathan Truong

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL MEDIA 최신 기사

그림 출처: Nathan Truong et al., arXiv:2608.18093, CC BY 4.0