ended6월 22일· 1 sources

Sparse KV Caches Cut Attention Scaling

Why it matters

Sparse key‑value caches collapse the quadratic blow‑up of softmax attention into a cost that grows near‑linearly with sequence length. By making each query attend to a tiny, top‑k subset of blockwise ...

1
Sources
+0
24h
Growth
11d
Active

Sources

Related Issues