ended9월 6일· 1 sources
SWE-Gate: Why Passing Tests Isn't Enough for Agent-Generated Code
Why it matters
Coding agents pass tests but fail code review. Repository-level benchmarks measure test passage but ignore review acceptance criteria. This is the blind spot in every benchmark from SWE-bench onward. ...
1
Sources
+0
24h
—
Growth
5d
Active