ended7월 30일· 1 sources
I gave the same fabricated answer to RAGAS and DeepEval. One scored it 0.0. The other scored it 1.0
Why it matters
Here's an output from a RAG system asserting a pricing claim it was never given, for a question its context couldn't answer. I ran it past the two most popular LLM-as-judge faithfulness metrics, five ...
1
Sources
+0
24h
—
Growth
3d
Active