new2시간 전· 1 sources

Qwen3.8-27B on One RTX 3090 vs Two: +20% Decode, +14% Cold Prefill, and 3x on Cached Prompts

Why it matters

TL;DR: Qwen3.8-27B — the dense 27.8B that dropped on August 14 — fits on a single RTX 3090 at W4A16 and decodes at ~125–155 tokens/sec. Splitting it across two cards with tensor parallelism buys rough...

1
Sources
+1
24h
Growth
1d
Active

Sources

Related Issues