ended5월 3일· 1 sources
AI Safety's Achilles' Heel: The Single Vector Controlling Model Refusal
LLM 안전망의 허점: 단 하나의 '거절 벡터'로 무너지는 AI 보안
Why it matters
This study exposes that AI refusal behavior is mediated by a surprisingly simple one-dimensional subspace, making safety filters easy to surgically disable. It underscores the fragility of current fine-tuning methods and opens new doors for precise mechanistic control over AI behavior.
1
Sources
+0
24h
—
Growth
133d
Active
Language ModelsLLM SafetyRefusal MechanismResidual StreamWhite-box Jailbreak