new1시간 전· 1 sources
Cutting 70% of RAG context tokens and keeping the answers identical (measured)
Why it matters
Your RAG pipeline retrieves 12 chunks because the retrieval score said "maybe". Your LLM reads all of them. You pay for all of them. And the answer quality was decided by chunks 2 and 7 anyway. On Sep...
1
Sources
+1
24h
—
Growth
1d
Active