rising4월 27일· 2 sources

OpenAI Moves Beyond SWE-bench as Frontier Models Outgrow Current Metrics

OpenAI, SWE-bench Verified 은퇴 선언... "현대 AI의 진화 속도 못 따라가"

Why it matters

OpenAI is halting its use of SWE-bench Verified because current frontier models have reached a performance level that exceeds the benchmark's ability to measure meaningful progress. This shift signals a critical need for more sophisticated evaluation frameworks that can accurately test AI's complex reasoning and long-context problem-solving in real-world software engineering.

2
Sources
+0
24h
Growth
147d
Active
AI BenchmarkingSWE-bench VerifiedFrontier ModelsData ContaminationSoftware EngineeringSWE-bench ProOpenAI

Sources

Related Issues