ended9월 7일· 1 sources
Speculative Decoding in vLLM on AMD GPUs
Why it matters
Exploring Speculative Decoding in vLLM on AMD GPUs TL;DR: Speculative decoding allows vLLM to verify multiple drafted tokens in a single target-model pass. In our experiments, its effect on output-tok...
1
Sources
+0
24h
—
Growth
14d
Active