Deployment & Production  ·  BraivIQ AI Engineering Playbook

Inside LLM Inference Serving: Prefill, Decode, KV Cache, Continuous Batching, Speculative Decoding And Disaggregation - Explained For Senior Engineers

Every senior engineer running AI in production pays for inference, but surprisingly few can explain what happens between a request arriving and tokens streaming back - and that layer decides your latency, your throughput and your bill. In 2026 the serving stack has matured into a recognisable end-state: vLLM or SGLang, paged attention with continuous batching, automatic prefix caching, an FP8 KV cache, speculative decoding with EAGLE-2, and prefill/decode disaggregation, with multi-tenant LoRA on top. Each of those is an answer to a specific physical problem in how transformers generate text. This educational deep-dive, for senior engineers and CTOs, walks the request path: why generation has two phases with opposite bottlenecks, why the KV cache is the resource everything fights over, how continuous batching and chunked prefill keep GPUs busy without starving anyone, when speculative decoding pays and when it does not, why the newest systems split prefill and decode onto different hardware, and how to reason about cost per token from first principles rather than from a price list.

 ·  14 min read  ·  By BraivIQ Engineering

Inside LLM Inference Serving: Prefill, Decode, KV Cache, Continuous Batching, Speculative Decoding And Disaggregation - Explained For Senior Engineers

Prefill vs decode - Two phases with opposite bottlenecks - compute-bound prompt processing, memory-bound token generation  ·  KV cache - The resource every request fights over; paged attention is what stopped it fragmenting GPU memory  ·  Continuous batching - Admit new requests mid-generation so the GPU never idles waiting for the slowest sequence  ·  2026 end-state - vLLM/SGLang + paged attention + prefix caching + FP8 KV + EAGLE-2 speculation + disaggregation + multi-tenant LoRA

Ask a room of senior engineers who run AI in production what they pay per million tokens and most will know. Ask them what physically happens on the GPU between a request arriving and tokens streaming back, and the room goes quiet - which is a problem, because that layer, the inference serving stack, is what decides your latency, your throughput and your bill, and it is now the subject of an unusually deep body of engineering. In 2026 the production stack has converged on a recognisable end-state: vLLM or SGLang as the engine, paged attention with continuous batching, automatic prefix caching, an FP8 KV cache, speculative decoding with EAGLE-2, and disaggregation of prefill and decode - increasingly on the same node - with multi-tenant LoRA serving layered on top. None of those is a buzzword; each is a precise answer to a physical constraint in how a transformer generates text, and understanding the constraint is what lets you choose the right knobs for your workload instead of copying someone else's configuration. As an AI Agency Developer London that runs self-hosted models for clients who cannot send data to an API, we think inference internals are the most under-taught topic in production AI, and this educational deep-dive is the request path, explained from first principles.

The KV Cache: The Resource Everything Fights Over

The state that decode reads at every step is the KV cache - the attention keys and values for every token processed so far, for every layer - and it is the central resource of inference serving, because it grows with sequence length, it must live in GPU memory, and every concurrent request has its own. On a large model a single long conversation's cache runs to gigabytes, so the number of requests a GPU can serve at once is bounded not by compute but by how many caches fit. Early servers allocated each request's cache as one contiguous block sized for the maximum possible length, which wasted most of it and fragmented memory so badly that GPUs ran at a fraction of their capacity. vLLM's paged attention borrowed the operating system's answer: allocate the cache in fixed-size pages on demand, indexed through a page table, so memory is used as sequences actually grow, freed blocks are reused, and fragmentation disappears - which alone multiplied achievable concurrency. Two refinements now sit on top. Automatic prefix caching recognises that many requests share a prefix - the same system prompt, the same retrieved document, the same few-shot examples - and reuses the already-computed pages for the shared portion rather than recomputing them, turning a long shared system prompt from a per-request prefill cost into a one-off. And FP8 KV cache stores the keys and values at eight-bit precision, halving the memory per token with negligible quality loss, so twice the concurrency fits. If you remember one thing about serving economics, remember that you are managing KV cache memory, and every technique is judged by what it does to it.

  • Concurrency is bounded by KV cache memory, not compute - how many sequences' state fits on the GPU decides throughput.
  • Paged attention allocates the cache in on-demand pages via a page table, eliminating the fragmentation that idled early servers.
  • Automatic prefix caching reuses computed pages for shared prefixes - system prompts, shared documents - so they are prefilled once.
  • FP8 KV cache halves memory per token with negligible quality loss, doubling the concurrency that fits.
  • Design prompts for the cache - put shared, stable content first and variable content last so prefixes actually match.

Keeping The GPU Busy: Continuous Batching And Chunked Prefill

Because decode is memory-bound, a GPU generating one token for one request is wasting almost all of its arithmetic; generating one token each for sixty requests in the same step costs little more, because the weights are read once and shared. Batching is therefore the foundation of throughput, but the naive form - assemble a batch, run it until every request finishes - is terrible, because sequences finish at different times and the batch idles waiting for the longest. Continuous batching, introduced by Orca and universal since, schedules at the granularity of a single step: at every decode iteration the scheduler evicts requests that finished, admits new requests waiting in the queue, and runs the next step on the current set, so the GPU is always working on a full batch and a new request's first token is not delayed until an old batch drains. The complication is prefill: a newly admitted request needs its prompt processed, and a long prompt's compute-bound prefill, dropped into a step full of memory-bound decodes, stalls every other request's next token for as long as it runs - visible to users as everyone's response freezing when someone pastes a large document. Chunked prefill is the fix: a long prompt's prefill is split into chunks that are interleaved with decode steps across several iterations, so prefill progresses without ever monopolising a step, and time-to-first-token for the new request is traded, slightly, for stable inter-token latency for everyone. Tuning the chunk size against your prompt-length distribution is one of the highest-leverage serving decisions, and it is workload-specific: RAG-heavy traffic with long contexts needs different settings from short chat turns.

Speculative Decoding, Disaggregation, And The Cost Per Token

Two further techniques address latency rather than throughput, and knowing which problem you have tells you whether to use them. Speculative decoding attacks the one-token-per-step limit of decode: a small, fast draft mechanism - EAGLE-2 being the current standard, which predicts several tokens ahead using the target model's own hidden state - proposes a handful of tokens, and the large model verifies them all in a single forward pass, accepting the prefix that matches what it would have produced. Because verification is one batched step rather than several sequential ones, a request's tokens arrive faster with output identical to the unaccelerated model. It shines at low concurrency, where per-request latency matters and the GPU has spare arithmetic to verify drafts; at high concurrency the GPU is already saturated by batching and speculation adds work for little gain, which is why it is the recommended tool for latency-sensitive, lightly-loaded workloads and not a universal switch. Disaggregation attacks the prefill-decode interference at its root: run prefill and decode on separate devices - or, in the newest designs, on separate partitions of the same GPU - so that prefill's compute-heavy bursts never stall decode's memory-bound steady stream, transferring the freshly-built KV cache from the prefill worker to the decode worker over a fast link. It buys stable decode latency and lets each phase run on hardware suited to it, at the cost of KV transfer and scheduling complexity, which frameworks like FlowKV and the same-node disaggregation in modern engines exist to manage. Put it all together and cost per token stops being a mystery: it is GPU-seconds per token, driven by how many concurrent sequences your KV cache fits (paged attention, prefix caching, FP8), how fully batched each step is (continuous batching, chunked prefill), and how many tokens each step yields (speculation) - which is why a well-tuned stack can be several times cheaper than a default one on identical hardware.

The Bottom Line

The inference serving stack decides your latency, throughput and bill, and every technique in the 2026 end-state - vLLM or SGLang with paged attention, continuous batching, chunked prefill, automatic prefix caching, FP8 KV cache, EAGLE-2 speculative decoding, prefill/decode disaggregation and multi-tenant LoRA - is a precise answer to a physical constraint in how transformers generate text. Generation has a compute-bound prefill and a memory-bound decode that fight over the GPU; the KV cache is the resource that bounds concurrency, which paged attention stopped fragmenting, prefix caching stopped recomputing and FP8 halved; continuous batching keeps every step full while chunked prefill stops long prompts freezing everyone; speculative decoding buys latency at low concurrency and little at high; and disaggregation separates the phases onto suited hardware for stable decode. Cost per token is GPU-seconds per token, driven by fitted concurrency, batch fullness and tokens per step - so a stack tuned to your prompt lengths, concurrency and latency budget can be several times cheaper than a default on the same hardware. Understanding the constraints is what lets you choose knobs rather than copy them, and tuning self-hosted serving to a client's real workload is exactly the engineering we do.

References & Further Reading

  • vLLM Blog - inside vLLM: anatomy of a high-throughput LLM inference system: https://vllm.ai/blog/2025-09-05-anatomy-of-vllm
  • Spheron - LLM serving optimization: continuous batching, PagedAttention and chunked prefill on H100 (2026): https://www.spheron.network/blog/llm-serving-optimization-continuous-batching-paged-attention/
  • Prompt20 - how modern LLM inference works: prefill, decode, KV cache and disaggregated inference (the 2026 production stack): https://blog.prompt20.com/posts/disaggregated-inference/
  • arXiv - FlowKV: a disaggregated inference framework with low-latency KV cache transfer and load-aware scheduling: https://arxiv.org/pdf/2504.03775
  • arXiv - Nexus: proactive intra-GPU disaggregation of prefill and decode in LLM serving: https://arxiv.org/pdf/2507.06608