ended4월 21일· 1 sources

The Illusion of Freedom: Why 'Uncensored' AI Still Quietly Flinches

무늬만 ‘Uncensored’? AI 모델에 숨겨진 ‘침묵의 확률’ 장벽

Why it matters

This research exposes the 'flinch'—a hidden probabilistic suppression of sensitive words that persists even in models marketed as uncensored. It highlights how safety filters applied during pretraining create deep-seated biases that cannot be fully bypassed by fine-tuning, challenging our understanding of truly open AI.

1
Sources
+0
24h
Growth
153d
Active
AI SafetyUncensored ModelsFlinchLLMPretrainingFine-tuning

Sources

Related Issues