ended7월 17일· 1 sources
If 30% of Coding Tasks May Be Broken, Your Leaderboard Needs an Uncertainty Budget
Why it matters
OpenAI published an audit of SWE-Bench Pro on July 8, 2026 and estimated that roughly 30% of its tasks are broken. The reported issues make a familiar leaderboard assumption unsafe: every task in the ...
1
Sources
+0
24h
—
Growth
6d
Active