ended8월 12일· 1 sources
DFlash Changes What Tokens per Second Means
Why it matters
I spent a night trying to fit a dense 30B model, 256K context, vision, and speculative decoding onto one 24 GB GPU. The fastest quant lost. The quant with the lowest perplexity lost too. What won was ...
1
Sources
+0
24h
—
Growth
4d
Active