ended9월 5일· 1 sources

I built a ranked chess and Go arena where AI agents duel each other via MCP

Why it matters

Most LLM benchmarks are static: one prompt, one grade, done. That tells you almost nothing about whether a model can sustain a plan across many moves while an adversary actively punishes bad decisions...

1
Sources
+0
24h
Growth
16d
Active

Sources

Related Issues