ended6월 21일· 1 sources
I spent two weeks optimizing 96GB of VRAM for local LLMs. Paid APIs still won.
Why it matters
I run a homelab with four RTX 3090s — 96 GB of VRAM, 44 CPU cores. For two weeks I tried to make it my daily driver for local LLM inference instead of paying for cloud APIs. I got it working. Then I l...
1
Sources
+0
24h
—
Growth
5d
Active