ended7월 30일· 1 sources

I gave the same fabricated answer to RAGAS and DeepEval. One scored it 0.0. The other scored it 1.0

Why it matters

Here's an output from a RAG system asserting a pricing claim it was never given, for a question its context couldn't answer. I ran it past the two most popular LLM-as-judge faithfulness metrics, five ...

1
Sources
+0
24h
Growth
3d
Active

Sources

Related Issues