ended4월 21일· 1 sources
Hardware Isn't the Bottleneck: Pushing Qwen3.5 to 207 tok/s on Consumer GPUs
"반도체 탓 하지 마라" RTX 3090으로 Qwen3.5-27B 속도 5배 끌어올린 소프트웨어의 힘
Why it matters
The Luce-Org project demonstrates that handcrafted CUDA kernels and advanced speculative decoding can extract massive performance from older consumer hardware. By bypassing standard inference bottlenecks, they've enabled large models like Qwen3.5-27B to run at speeds previously reserved for enterprise-grade clusters, proving that software optimization remains the ultimate frontier for LLM accessibility.
1
Sources
+0
24h
—
Growth
153d
Active
Qwen3.5RTX 3090Speculative DecodingDFlashMegakernelLLM Inference