new9시간 전· 1 sources
We Thought the LLM Was Wrong. Our Safety Detector Was Wrong.
Why it matters
There is a hidden dependency in a lot of LLM safety benchmarks: the detector. You send an adversarial prompt to a model, collect its response, and then some classifier decides whether that response re...
1
Sources
+1
24h
—
Growth
1d
Active