ended8월 12일· 1 sources

DFlash Changes What Tokens per Second Means

Why it matters

I spent a night trying to fit a dense 30B model, 256K context, vision, and speculative decoding onto one 24 GB GPU. The fastest quant lost. The quant with the lowest perplexity lost too. What won was ...

1
Sources
+0
24h
Growth
4d
Active

Sources

Related Issues