ended7월 3일· 1 sources

The hard part of attacking an AI isn't breaking it. It's telling real harm from fake.

Why it matters

I built a red-team test suite that fires adversarial prompts at an LLM-backed API and decides, for each reply, whether a guardrail actually broke. It is the project where I stopped writing tests that ...

1
Sources
+0
24h
Growth
80d
Active

Sources

Related Issues