ended7월 9일· 1 sources
LLM-as-judge disagrees with itself between runs
Why it matters
The flap I had a faithfulness gate on merge: judge scores every case, the mean has to clear 0.80. One Tuesday it failed at 0.79. I re-ran the identical job, no code change, no prompt change, and it pa...
1
Sources
+0
24h
—
Growth
74d
Active