ended4월 16일· 1 sources
AI 에이전트 스킬, 벤치마크 성능의 절반도 현실에서 안 나온다
Why it matters
Research from UC Santa Barbara and MIT reveals that AI agents' skill utilization underperforms dramatically in real-world scenarios compared to benchmarks—Claude Opus 4.6 drops from 55.4% to 40.1% accuracy when skills must be discovered and selected rather than provided. Poor skill selection and retrieval accuracy emerge as critical bottlenecks, with weaker models sometimes performing worse when using skills. This challenges the reliability of current benchmark-based performance claims in practical AI agent deployment.
1
Sources
+0
24h
—
Growth
158d
Active
AI agentsskill selectionbenchmark gapClaude Opusretrieval accuracy