ended3월 26일· 1 sources
GPT-5, Claude, Gemini All Score Below 1% - ARC AGI 3 Just Broke Every Frontier Model
GPT-5, Claude, Gemini 모두 1% 미만 — ARC-AGI-3가 모든 최전선 모델을 무너뜨리다
Why it matters
ARC-AGI-3, launched on March 25, 2026, completely overhauls the ARC benchmark by replacing static grid puzzles with interactive, game-like environments where AI agents must discover rules and solve problems without any instructions. Frontier LLMs like GPT-5 and Claude score below 1%, while simple CNN and graph-search methods reach 12.58%, highlighting a massive gap in interactive reasoning capabilities. The competition offers over $2 million in prizes and tests exploration, world modeling, goal-setting, and strategic planning — capabilities where current large language models fundamentally struggle.
1
Sources
+0
24h
—
Growth
171d
Active
ARC-AGI-3GPT-5ClaudeGeminiFrançois CholletKaggle