LLM inference & serving
PageServe
A vLLM-style inference engine, grown from a naive generate() loop in 11 phases.
Continuous batching, a paged KV cache over a physical tensor pool, prefill/decode budgets, and CPU swap, so under load it slows down politely instead of OOMing. Along the way I found a cache bug quietly charging 1.1 ms per token, and evicted it.
- −98.5%
- mean TTFT at 4-way load
1,418 → 20.9 ms - −26.5%
- wall-clock latency
vs sequential - 11
- build phases,
each runnable