ended8월 1일· 1 sources
Why INT4 Weight-Only Quantization Doesn't Speed Up Prefill
Why it matters
You benchmark a 70B model with batch_size=1 , one prompt, one stream. FP16 gives you 18 tokens/sec. You swap in an AWQ INT4 checkpoint and get 55 tokens/sec. Three times faster, same GPU, ~1 point of ...
1
Sources
+0
24h
—
Growth
4d
Active