new3시간 전· 1 sources

KV cache cut by ~45% with near‑same accuracy

Why it matters

Grouped Value Attention slashes transformer KV memory by roughly 45 % without hurting benchmark scores. By storing only grouped values and reconstructing keys on the fly, it eliminates the need to mat...

1
Sources
+1
24h
—
Growth
1d
Active

Sources

Related Issues