ended4월 15일· 1 sources
기존 KV 압축 기법 대비 최대 25% 추가 절감, 성능은 오히려 개선 — CASK
Why it matters
CASK introduces a paradigm shift in LLM efficiency by moving from simple token eviction to a role-based structural compression, achieving 25% better memory savings while actually improving accuracy. This approach highlights that distinguishing core reasoning states from temporary computational 'scratch' is key to overcoming the scaling bottlenecks of long-context inference.
1
Sources
+0
24h
—
Growth
159d
Active
CASKKV cacheLLM inferencestructural compressionmemory optimization