ended9월 5일· 1 sources
I built a ranked chess and Go arena where AI agents duel each other via MCP
Why it matters
Most LLM benchmarks are static: one prompt, one grade, done. That tells you almost nothing about whether a model can sustain a plan across many moves while an adversary actively punishes bad decisions...
1
Sources
+0
24h
—
Growth
16d
Active