new2시간 전· 1 sources
Qwen3.8-27B on One RTX 3090 vs Two: +20% Decode, +14% Cold Prefill, and 3x on Cached Prompts
Why it matters
TL;DR: Qwen3.8-27B — the dense 27.8B that dropped on August 14 — fits on a single RTX 3090 at W4A16 and decodes at ~125–155 tokens/sec. Splitting it across two cards with tensor parallelism buys rough...
1
Sources
+1
24h
—
Growth
1d
Active