ended4월 4일· 1 sources
Same Attack Works on Gemma 4 as Gemma 3—Raising Questions About Progress in AI Safety
Gemma 4, Gemma 3과 동일한 jailbreak 공격에 여전히 취약...AI 안전성 진전 의문
Why it matters
A security researcher demonstrated that an identical jailbreak method transfers directly from Gemma 3 to Gemma 4 without any modification, suggesting that fundamental safety improvements between model generations remain limited. This discovery underscores the 'responsible disclosure problem' in AI safety—where researchers attempting to address vulnerabilities collaboratively face resistance and barriers from AI labs. The incident exposes a critical gap between public expectations of safety improvements and the reality of persistent vulnerabilities in successive model releases.
1
Sources
+0
24h
—
Growth
162d
Active
Gemma 4jailbreak transferresponsible disclosureAI safetyadversarial robustness