ended7월 9일· 1 sources

LLM-as-judge disagrees with itself between runs

Why it matters

The flap I had a faithfulness gate on merge: judge scores every case, the mean has to clear 0.80. One Tuesday it failed at 0.79. I re-ran the identical job, no code change, no prompt change, and it pa...

1
Sources
+0
24h
Growth
74d
Active

Sources

Related Issues