ended7월 23일· 1 sources
Can a MUD evaluate LLMs? A $99 proof of concept
Why it matters
Old worlds for new agents. CrucibleBench places language models in a persistent MUD, a text world where NPCs remember, trust accumulates, and mistakes leave traces, and scores what they do over 50 tur...
1
Sources
+0
24h
—
Growth
60d
Active