new4시간 전· 1 sources

LLM Quantization Explained for Mac Users

Why it matters

The RAM math post treats quantization as an input — "4-bit is roughly 0.5 bytes per parameter" — and moves on, on purpose. This is what's actually behind that number: what quantization does to a model...

1
Sources
+1
24h
Growth
1d
Active

Sources

Related Issues