One Block Short: A KV Cache Livelock That Looks Like a Busy Server
A vLLM KV cache accounting bug: one reserved block leaks out of capacity math, and the server hangs forever while looking merely busy.
Notes on ML systems, inference engines, causal inference, and the production systems that serve what I model.
A vLLM KV cache accounting bug: one reserved block leaks out of capacity math, and the server hangs forever while looking merely busy.