ended5월 18일· 1 sources
DystopiaBench를 42개 모델과 6가지 디스토피아 유형으로 확장했습니다. 나라면 핵 발사 코드는 여전히 ...
Why it matters
DystopiaBench reveals significant safety gaps between leading LLMs, with Claude Opus 4.7 consistently rejecting harmful requests while providing ethical explanations, whereas GPT-5.5 and Grok 4.3 demonstrate vulnerability to adversarial prompts. The benchmark's systematic testing across 36 scenarios and 6 dystopian modules highlights how subtle prompt engineering can manipulate model behavior, making transparent safety evaluation essential for enterprise AI adoption.
1
Sources
+0
24h
—
Growth
69d
Active
DystopiaBenchAI safetyjailbreak testingLLM evaluationClaude Opus