ended4월 15일· 1 sources

기존 KV 압축 기법 대비 최대 25% 추가 절감, 성능은 오히려 개선 — CASK

Why it matters

CASK introduces a paradigm shift in LLM efficiency by moving from simple token eviction to a role-based structural compression, achieving 25% better memory savings while actually improving accuracy. This approach highlights that distinguishing core reasoning states from temporary computational 'scratch' is key to overcoming the scaling bottlenecks of long-context inference.

1
Sources
+0
24h
Growth
159d
Active
CASKKV cacheLLM inferencestructural compressionmemory optimization

Sources

Related Issues