ended4월 21일· 1 sources
The Illusion of Freedom: Why 'Uncensored' AI Still Quietly Flinches
무늬만 ‘Uncensored’? AI 모델에 숨겨진 ‘침묵의 확률’ 장벽
Why it matters
This research exposes the 'flinch'—a hidden probabilistic suppression of sensitive words that persists even in models marketed as uncensored. It highlights how safety filters applied during pretraining create deep-seated biases that cannot be fully bypassed by fine-tuning, challenging our understanding of truly open AI.
1
Sources
+0
24h
—
Growth
153d
Active
AI SafetyUncensored ModelsFlinchLLMPretrainingFine-tuning