ended3월 24일· 1 sources
AI Research Monthly: Feb-Mar 2026 — 21 Findings With Hard Data (The Comprehensive Edition)
AI 연구 월간 리포트: 2026년 2~3월 — 실측 데이터 기반 21가지 주요 발견 (종합판)
Why it matters
OpenAI's audit found 59.4% of SWE-bench Verified's hardest problems had flawed test cases and evidence of benchmark contamination, leading to inflated scores; the new SWE-bench Pro caps real performance at 57%. FeatureBench (ICLR 2026) revealed a 63-point gap between AI bug-fixing (74%) and feature-building (11%), showing these are fundamentally different skills. The piece also references Humanity's Last Exam improving from single digits to 37% within a year.
1
Sources
+0
24h
—
Growth
181d
Active
SWE-benchFeatureBenchbenchmark contaminationcoding AIHumanity's Last Exam