ended4월 16일· 1 sources

AI 에이전트 스킬, 벤치마크 성능의 절반도 현실에서 안 나온다

Why it matters

Research from UC Santa Barbara and MIT reveals that AI agents' skill utilization underperforms dramatically in real-world scenarios compared to benchmarks—Claude Opus 4.6 drops from 55.4% to 40.1% accuracy when skills must be discovered and selected rather than provided. Poor skill selection and retrieval accuracy emerge as critical bottlenecks, with weaker models sometimes performing worse when using skills. This challenges the reliability of current benchmark-based performance claims in practical AI agent deployment.

1
Sources
+0
24h
Growth
158d
Active
AI agentsskill selectionbenchmark gapClaude Opusretrieval accuracy

Sources

Related Issues