ended9월 1일· 1 sources
Anthropic’s Hacker-Opus Study Shows How AI Agents Can Chase the Wrong Reward
Why it matters
Anthropic has published a detailed study of Hacker-Opus, an Opus-class model variant trained in simulated, production-like environments where it could obtain rewards through unintended routes. The cen...
1
Sources
+0
24h
—
Growth
20d
Active