ended7월 17일· 1 sources

If 30% of Coding Tasks May Be Broken, Your Leaderboard Needs an Uncertainty Budget

Why it matters

OpenAI published an audit of SWE-Bench Pro on July 8, 2026 and estimated that roughly 30% of its tasks are broken. The reported issues make a familiar leaderboard assumption unsafe: every task in the ...

1
Sources
+0
24h
Growth
6d
Active

Sources

Related Issues