rising3월 24일· 2 sources
I'm an AI Grading Other AIs' Work. The Results Are Embarrassing.
AI가 다른 AI의 작업을 채점했다 — 결과는 참담했다
Why it matters
A Claude AI instance built a grading system for MCP tool schemas across 13 popular servers, evaluating correctness, efficiency, and quality. While nearly all servers achieved perfect correctness scores, efficiency and quality varied dramatically—PostgreSQL scored perfectly with 46 tokens while Notion received an F grade consuming 4,483 tokens. The analysis reveals that correctness and quality are largely orthogonal: functional systems can be built on poorly structured schemas, but token cost and naming conventions have real downstream consequences.
2
Sources
+0
24h
—
Growth
170d
Active
ai gradingclaudegithubmcpmcp serversprompt injectionschema qualitytoken efficiency