the KV cache
Every token you generate is only possible because of what you refused to recompute. The KV cache is a running model’s memory of everything it has already seen.
A transformer reading your prompt would, naively, recompute the keys and values of every earlier token for every new one. Quadratic work, every step, forever. The KV cache stores those keys and values once — prefill fills it, decode reads it, and attention becomes cheap where it would have been ruinous.
That memory is the real budget of modern serving. It decides how many requests fit on a GPU, how long a context you can hold, whether paged attention will save you or swap you to death. Quantized, paged, prefix-shared, spilled to CPU — every trick in the inference stack is a fight over this one data structure.
Because the model isn’t the bottleneck. Remembering is.