ended8월 10일· 1 sources
Self-hosting a lite agent backend on one TPU: Gemma 4 E2B + vLLM on a v5e-1
Why it matters
Self-hosting a lite agent backend on one TPU chip A single Google Cloud TPU v5e chip — 16 GB of HBM, about $0.58/hour on spot — will serve google/gemma-4-E2B-it under vLLM at 1,496 output tokens/sec a...
1
Sources
+0
24h
—
Growth
42d
Active