ended7월 26일· 1 sources
Is Speculative Decoding's Speedup a Hardware Problem or a Model Problem?
Why it matters
A follow-up/sub-part to Part 3 of the LLM inference internals series. Part 3 built sampling-mode speculative decoding with KV caching on both the draft and verifier sides, and it worked correctly, but...
1
Sources
+0
24h
—
Growth
6d
Active