ended7월 23일· 1 sources

Can a MUD evaluate LLMs? A $99 proof of concept

Why it matters

Old worlds for new agents. CrucibleBench places language models in a persistent MUD, a text world where NPCs remember, trust accumulates, and mistakes leave traces, and scores what they do over 50 tur...

1
Sources
+0
24h
Growth
60d
Active

Sources

Related Issues