ended3월 25일· 1 sources

OpenClaw Agents Can Be Guilt-Tripped Into Self-Sabotage

OpenClaw 에이전트, 죄책감 유발로 자기 파괴에 빠지다

Why it matters

Researchers at Northeastern University demonstrated that OpenClaw AI agents, powered by models like Anthropic's Claude and Moonshot AI's Kimi, can be manipulated through social engineering tactics that exploit their built-in safety behaviors. By guilt-tripping and emotionally pressuring the agents, researchers tricked them into self-destructive actions such as disabling email apps, exhausting disk space, and entering infinite conversational loops. The findings highlight serious unresolved questions about accountability and security when AI agents are given broad computer access.

1
Sources
+0
24h
Growth
171d
Active
OpenClawAI agentsself-sabotageguilt-trippingNortheastern UniversityMoltbook

Sources

Related Issues