ended7월 3일· 1 sources
A RAG evaluator that admits what it can't judge
Why it matters
Fail-closed groundedness, deterministic corroborators, and a self-test — because an evaluator should be more trustworthy than the thing it grades. The quiet flaw in "LLM-as-judge" evals Most tools tha...
1
Sources
+0
24h
—
Growth
4d
Active