ended3월 19일· 1 sources

EsoLang-Bench: Evaluating Genuine Reasoning in LLMs via Esoteric Languages

EsoLang-Bench: 난해한 프로그래밍 언어를 활용한 LLM의 진정한 추론 능력 평가

Why it matters

EsoLang-Bench is a new benchmark evaluating LLM code generation across five esoteric programming languages where training data is 5,000–100,000x scarcer than Python. Frontier models that score 85–95% on standard benchmarks collapse to 0–11% accuracy on equivalent esoteric tasks, with Whitespace remaining completely unsolved at 0%. The results demonstrate that high scores on mainstream language benchmarks largely reflect memorization of training data rather than genuine reasoning or general programming ability.

1
Sources
+0
24h
Growth
180d
Active
EsoLang-BenchLLMesoteric languagescode generationBrainfuckWhitespace

Sources

Related Issues