ended6월 22일· 1 sources
Sparse KV Caches Cut Attention Scaling
Why it matters
Sparse key‑value caches collapse the quadratic blow‑up of softmax attention into a cost that grows near‑linearly with sequence length. By making each query attend to a tiny, top‑k subset of blockwise ...
1
Sources
+0
24h
—
Growth
11d
Active