ended5월 14일· 1 sources

LLMs Lose Performance After Launch. New Data Exposes the Pattern.

LLM, 출시 후 성능이 점점 떨어진다... 데이터가 증명했다

Why it matters

An independent researcher has published an interactive visualization revealing a hidden industry pattern: large language models systematically decline in performance weeks or months after their release. By tracking daily ELO ratings from the public LM Arena leaderboard, this tool shows that headline-grabbing benchmark announcements don't reflect real-world performance decay, likely caused by aggressive quantization or additional filters applied post-launch. For users and organizations evaluating LLMs, this underscores that launch-day metrics are unreliable predictors of sustained quality.

1
Sources
+0
24h
Growth
4d
Active
model degradationArena ELOLM ArenaELO ratingperformance decaybenchmark

Sources

Related Issues