ended5월 24일· 1 sources
Qwen 3.6 27B and 35B MTP vs Standard on 16GB GPU
Why it matters
I tested Speculative decoding (Multi-Token Prediction, MTP) performance in Qwen 3.6 27B and 35B on an RTX 4080 with 16 GB VRAM. For a broader view of token speeds and VRAM trade-offs across more model...
1
Sources
+0
24h
—
Growth
3d
Active