ended7월 28일· 1 sources

KV Cache Quantization: I Stretched Qwen 35B's Context 8 on 12GB VRAM

Why it matters

600 MiB of headroom My RTX 4070 was running Qwen 35B beautifully after the --cpu-moe trick from a previous run. The tokens/sec were where I wanted them. VRAM sat at 11,714 MiB out of 12,281 — 95% full...

1
Sources
+0
24h
Growth
4d
Active

Sources

Related Issues