new4시간 전· 1 sources
LLM Quantization Explained for Mac Users
Why it matters
The RAM math post treats quantization as an input — "4-bit is roughly 0.5 bytes per parameter" — and moves on, on purpose. This is what's actually behind that number: what quantization does to a model...
1
Sources
+1
24h
—
Growth
1d
Active