ended7월 28일· 1 sources
KV Cache Quantization: I Stretched Qwen 35B's Context 8 on 12GB VRAM
Why it matters
600 MiB of headroom My RTX 4070 was running Qwen 35B beautifully after the --cpu-moe trick from a previous run. The tokens/sec were where I wanted them. VRAM sat at 11,714 MiB out of 12,281 — 95% full...
1
Sources
+0
24h
—
Growth
4d
Active