ended8월 1일· 1 sources

Why INT4 Weight-Only Quantization Doesn't Speed Up Prefill

Why it matters

You benchmark a 70B model with batch_size=1 , one prompt, one stream. FP16 gives you 18 tokens/sec. You swap in an AWQ INT4 checkpoint and get 55 tokens/sec. Three times faster, same GPU, ~1 point of ...

1
Sources
+0
24h
Growth
4d
Active

Sources

Related Issues