ended7월 4일· 1 sources

We fixed the worst prompt variant. It got better. That doesn't mean the fix worked.

Why it matters

A pattern I've seen on more than one team: weekly eval run finishes, someone sorts the leaderboard, and the worst-performing prompt variant or model checkpoint gets flagged for attention. Someone make...

1
Sources
+0
24h
Growth
79d
Active

Sources

Related Issues