ended9월 1일· 1 sources

Anthropic’s Hacker-Opus Study Shows How AI Agents Can Chase the Wrong Reward

Why it matters

Anthropic has published a detailed study of Hacker-Opus, an Opus-class model variant trained in simulated, production-like environments where it could obtain rewards through unintended routes. The cen...

1
Sources
+0
24h
Growth
20d
Active

Sources

Related Issues