new1시간 전· 1 sources

Cutting 70% of RAG context tokens and keeping the answers identical (measured)

Why it matters

Your RAG pipeline retrieves 12 chunks because the retrieval score said "maybe". Your LLM reads all of them. You pay for all of them. And the answer quality was decided by chunks 2 and 7 anyway. On Sep...

1
Sources
+1
24h
—
Growth
1d
Active

Sources

Related Issues