ended9월 6일· 1 sources

SWE-Gate: Why Passing Tests Isn't Enough for Agent-Generated Code

Why it matters

Coding agents pass tests but fail code review. Repository-level benchmarks measure test passage but ignore review acceptance criteria. This is the blind spot in every benchmark from SWE-bench onward. ...

1
Sources
+0
24h
Growth
5d
Active

Sources

Related Issues