ended8월 6일· 1 sources

One Long Prompt Shouldn't Freeze Everyone's Tokens: Prefill/Decode Disaggregation

Why it matters

An LLM request is two workloads in a trench coat — a heavy, bursty prefill and a stream of tiny latency-sensitive decodes. Running them on the same engines lets one big prompt stall everyone. Splittin...

1
Sources
+0
24h
Growth
46d
Active

Sources

Related Issues