ended7월 3일· 1 sources
The hard part of attacking an AI isn't breaking it. It's telling real harm from fake.
Why it matters
I built a red-team test suite that fires adversarial prompts at an LLM-backed API and decides, for each reply, whether a guardrail actually broke. It is the project where I stopped writing tests that ...
1
Sources
+0
24h
—
Growth
80d
Active