ended9월 7일· 1 sources

Speculative Decoding in vLLM on AMD GPUs

Why it matters

Exploring Speculative Decoding in vLLM on AMD GPUs TL;DR: Speculative decoding allows vLLM to verify multiple drafted tokens in a single target-model pass. In our experiments, its effect on output-tok...

1
Sources
+0
24h
Growth
14d
Active

Sources

Related Issues