Back

Serving LLMs: What Actually Determines Your Throughput and Latency

Aug 25, 2026 (2w ago)

Serving LLMs: What Actually Determines Your Throughput and Latency

Most teams get the model working and then get blindsided by the serving bill. The demo felt instant; production feels sluggish under ten concurrent users. Nothing about the model changed — what changed is that inference is two completely different workloads wearing one API, and almost every surprise traces back to that.

This post is about the mechanics: what the hardware is actually doing, which knob moves which metric, and which trade-offs are unavoidable.

The two phases, and why they behave nothing alike

Generation runs in two phases with opposite bottlenecks.

Prefill processes your entire prompt in one shot. Every token attends in parallel, the GPU gets big dense matrix multiplies, and utilization is high. Prefill is compute-bound. Cost scales with prompt length. That one pass produces the first token and fills the KV cache — time-to-first-token ends there.

Decode produces one token at a time. Each step reads the whole weight matrix out of HBM to compute a single token's worth of math. Arithmetic intensity is terrible. Decode is memory-bandwidth-bound. Cost scales with tokens generated, and it barely cares how much compute the card has.

Consequences that catch people out:

  • A long prompt with a short answer is a prefill-dominated request. A short prompt with a long answer is decode-dominated. These need different tuning and they queue against each other.
  • Batching helps decode enormously (you amortize one weight read across many sequences) and helps prefill much less (it was already compute-saturated).
  • A faster card with more FLOPs but the same bandwidth moves prefill and does close to nothing for single-stream decode.

The metrics, and how they fight each other

  • TTFT (time to first token): queue wait + prefill. What the user perceives as "did it hang?"
  • ITL (inter-token latency): the gap between streamed tokens. Perceived reading speed.
  • End-to-end latency: TTFT + ITL × tokens generated. The only metric that matters for non-streaming calls.
  • Throughput: two different numbers — output tokens/sec across the whole server, and requests/sec completed. Cost per request lives here.

The core tension: throughput and per-request latency are bought from the same budget. Larger batches raise aggregate tokens/sec while lowering each individual stream's token rate, because more sequences share each weight read. Prioritizing an incoming prefill delays every decode already in flight, so admitting requests aggressively improves TTFT for newcomers and worsens ITL for everyone mid-generation.

Two things worth being precise about:

  • "Tokens per second" is never one number. Prefill tokens/sec and decode tokens/sec are different regimes measured in the same unit. Reporting a single blended figure hides which one you actually improved.
  • A throughput win can read as a latency regression to a single user. That's not a bug in your change; it's the trade you made. Decide beforehand which one you're optimizing.

If you're serving multi-step agent loops, note that a loop multiplies request count and makes TTFT compound across every step — ITL matters more than it looks when a "single" user action is really eight sequential calls.

Measure it from the client, not the server

Server-side numbers omit queue time, network, and the tokenizer. Probe the endpoint the way a user hits it: a streaming POST to /v1/chat/completions, one perf_counter() stamp per chunk that actually carries content (skip empty deltas and the [DONE] sentinel), and the rest is arithmetic on those stamps.

e2e          = time.perf_counter() - t0             # once the stream has closed
ttft         = stamps[0] - t0
gaps         = [b - a for a, b in zip(stamps, stamps[1:])]
itl_median   = statistics.median(gaps)
itl_p95      = sorted(gaps)[int(0.95 * len(gaps))]  # the one that bites
decode_tok_s = (len(stamps) - 1) / sum(gaps)

Run this at concurrency 1, then at your target concurrency. The delta between the two is your real capacity story — one stream tells you the latency floor, many streams tell you where the batch starts eating it.

The KV cache is your actual constraint

Attention needs the keys and values of every previous token. Recomputing them each step would be quadratic, so you cache them. That cache is the thing that fills your GPU.

Per token, per sequence:

kv_bytes_per_token = 2 (K and V) × num_layers × num_kv_heads × head_dim × bytes_per_element
dtype bytes_per_element
fp32 4
fp16, bf16 2
fp8, int8 1
int4 0.5

Note num_kv_heads, not the attention head count. Grouped-query and multi-query attention share KV heads across query heads specifically to shrink this term — using the query head count silently overstates KV usage by the grouping factor on any modern model.

The important property: this grows with batch size × sequence length, both of which are runtime variables. Weights are a fixed cost you pay once at load. KV is a variable cost that scales with traffic, and it is usually what actually runs you out of memory.

Take a hypothetical shape, chosen for round arithmetic and not a measurement of any model: 32 layers, 8 KV heads, head_dim 128, 7B parameters, weights and KV both fp16, max_seq_len 8,192, concurrency 32. Weights cost num_params × the weight dtype's bytes per element; KV costs the formula above × max_seq_len × concurrency; and a real budget adds roughly 10% on top of the two combined for activations, CUDA graphs, allocator fragmentation, and runtime.

Working that example by hand, because the shape of the answer is the point:

  • 2 × 32 × 8 × 128 × 2 bytes = 131,072 bytes = 128 KiB per token
  • × 8,192 tokens = exactly 1 GiB of KV cache per fully-extended request
  • × 32 concurrent requests = 32 GiB of KV, against ~13 GiB of fp16 weights

The variable part is more than twice the fixed part, and that's with grouped-query attention. If that same shape used full multi-head attention with 32 KV heads, the per-request figure would be 4 GiB instead of 1. Halving max_seq_len halves it too — which is why your advertised context window is a capacity decision, not a marketing one.

Paged attention, and why naive allocation wastes so much

The obvious implementation reserves one contiguous buffer per sequence, sized for the maximum length it might reach. Two failures follow: internal fragmentation, because most requests stop far short of the max and the reservation stays pinned; and external fragmentation, because freed slots of mismatched sizes leave holes nothing fits into.

Paged attention borrows virtual memory. KV is stored in fixed-size blocks, and each sequence keeps a block table mapping logical positions to physical blocks — which need not be adjacent.

You allocate as sequences grow, waste at most one partial block each, and get copy-on-write sharing for free: parallel samples from one prompt, or many requests behind one system prompt, can point at the same physical blocks. Higher achievable concurrency at identical memory is where most of the "we switched engines and throughput jumped" stories actually come from.

Batching: the biggest single win

Strategy How it schedules Effect on throughput Effect on latency Use when
Static Fixed batch assembled offline, run to completion Good if inputs are uniform Irrelevant — no live users Offline/batch jobs
Dynamic Wait a short window, group whatever arrived, run to completion Better than one-at-a-time Adds queue delay; batch runs until its longest member finishes Uniform, short generations
Continuous (in-flight) Schedules per decode step; finished sequences leave, new ones join mid-flight Best by a wide margin New requests start without waiting for the batch to drain Basically all interactive serving

Dynamic batching's flaw is head-of-line blocking at the sequence level. Generation lengths vary wildly; a batch of eight where seven finish in 20 tokens and one runs to 2,000 keeps seven slots idle for the remaining 1,980 steps.

Continuous batching schedules at token granularity instead. Every decode step, the scheduler rebuilds the running set: completed sequences release their KV blocks immediately, waiting requests are admitted into the freed capacity. GPU occupancy stays high instead of decaying as a batch drains. For a decode-bound workload — which is most chat traffic — this is the single highest-leverage change available, and it's the default in every serious serving engine. If you're running a hand-rolled loop over model.generate() behind a queue, this is the thing to fix first.

Quantization: what you buy and what you risk

Decode is bandwidth-bound, so bytes moved per token is the currency. Fewer bytes per weight means less to read each step.

Approach What shrinks What you gain What it costs
Weight-only (int8 / int4) Weights on disk and in HBM Large memory saving, faster bandwidth-bound decode; frees room for more KV Dequantization overhead per step; low-bit needs calibration; quality drift on hard tasks
Weight + activation (int8 / fp8) Weights and the activations flowing through Above, plus genuinely cheaper matmuls when the hardware has native low-precision units Activation outliers are the hard part; more accuracy risk; needs hardware support to pay off
KV cache quantization The KV cache itself Directly raises concurrency and usable context — attacks the real constraint Degradation shows up on long contexts, exactly where you needed it

Three things people get wrong:

  • Quantization is not automatically a speedup. It's reliably a memory win. On bandwidth-bound decode, less memory traffic converts to speed. On compute-bound prefill, weight-only quantization can be neutral or slightly negative, since you pay to dequantize before the same matmul runs.
  • The second-order win is often bigger than the first. Smaller weights leave more HBM for KV, which raises concurrency, which raises throughput. That path frequently beats the direct bandwidth gain.
  • Quality degradation is uneven. It hides on short generic prompts and appears on long context, structured output, and multi-step reasoning. Evaluate on your own traffic before and after — never assume a published quality claim transfers to your workload.

The rest of the toolbox

Speculative decoding. A cheap draft model proposes several tokens; the target model verifies them in one parallel pass, keeping the longest correct prefix. Output is identical to greedy decoding from the target — this is a latency optimization, not an approximation. It converts spare compute into fewer sequential steps, so it shines at low batch sizes where you're bandwidth-starved. Once you're batched heavily enough to be compute-bound, verification competes with real work and it can hurt aggregate throughput. Same shape as the quantization point: know which regime you're in.

Prefix caching. If many requests share a leading prefix — a fixed system prompt, a long tool schema, a retrieved context reused across turns of one conversation — its KV blocks can be computed once and reused, skipping that part of prefill entirely. Pure TTFT win, and it grows with prefix length and hit rate. Keep the shared portion at the very front and byte-identical; a per-request timestamp or user ID prepended to the system prompt destroys every hit.

Chunked prefill. A long prompt's prefill is one big step that stalls every in-flight decode, spiking ITL for other users. Splitting it into chunks interleaved with decode steps trades a little TTFT for that request against much steadier ITL for everyone else. This is the fix when one user's 30k-token paste makes the whole service stutter.

Tensor vs. pipeline parallelism. Both are for when weights plus working KV won't fit on one device — you don't reach for them voluntarily.

  • Tensor parallel splits individual matrices across devices; every layer requires a collective communication. Fast, but it wants high-bandwidth interconnect within a node.
  • Pipeline parallel splits layers across devices; communication is small (activations at stage boundaries) but a single request traverses stages sequentially, so you need multiple in-flight requests to keep stages busy.

Rule of thumb: tensor parallel inside a node, pipeline parallel across nodes, and prefer neither if a smaller or quantized model meets your quality bar on one device.

Capacity planning

Budget memory in this order:

  1. Weights — parameters × bytes per element. Fixed.
  2. KV cache — the estimator above, at your chosen max_seq_len and target concurrency. Variable.
  3. Activations — transient per-step working memory; scales with batch and hidden size.
  4. Overhead — runtime, CUDA graphs, allocator fragmentation, the framework itself. Leave real headroom; this is not zero.

Then work backwards. Fix the device budget, subtract weights and overhead, divide the remainder by KV bytes per token, and that quotient is your total token budget across all concurrent sequences. Every serving decision spends from it: max_seq_len × concurrency cannot exceed it. Serving a large advertised context to many users at once is arithmetically impossible on a fixed budget — pick which one you're actually selling.

Admission control. When the token budget is exhausted, you have three options: queue, preempt, or reject. Queueing inflates TTFT invisibly. Preemption (evicting a sequence's KV and recomputing or swapping it later) keeps the system live but wastes work. Rejecting is honest and protects everyone already in flight. Set an explicit queue depth and a maximum wait, and shed load past it — an unbounded queue converts a throughput problem into a timeout cascade. Separate queues for interactive and batch traffic keep a bulk job from starving chat.

Autoscaling. GPU cold starts are dominated by pulling a large image and loading weights, and they are slow enough that reactive scaling on a latency alarm arrives after the incident. Practical mitigations: bake weights into the image or stage them on fast local storage, keep a warm buffer replica, scale on a leading signal (queue depth or admitted tokens in flight, not GPU utilization — a busy decode batch shows low utilization by design), and scale in far more conservatively than you scale out.

A deployment shape

# Local shape check. Pin an explicit image digest in real use.
docker run --gpus all -p 8000:8000 \
  -v "$HOME/.cache/huggingface:/root/.cache/huggingface" \
  vllm/vllm-openai:PINNED_TAG \
  --model "$MODEL_ID" \
  --max-model-len 8192 \
  --max-num-seqs 64 \
  --enable-prefix-caching \
  --tensor-parallel-size 1

Wrapped in a Deployment, the fields worth arguing about:

Field Value here
replicas 2
image pinned by digest — your-registry/llm-serve@sha256:REPLACE_ME
resources.limits nvidia.com/gpu: 1
terminationGracePeriodSeconds 300
readinessProbe /health on port 8000, initialDelaySeconds: 60, periodSeconds: 10, failureThreshold: 30
volumeMounts / volumes hostPath: /mnt/local-ssd/models mounted at /models

The model id lives in env as MODEL_ID and is referenced from args as --model=$(MODEL_ID) — $(VAR) in args only expands for variables declared in env.

The load-bearing details: a readiness probe patient enough to cover weight loading, a termination grace period long enough that a rolling update doesn't guillotine active streams, and model files on fast local storage so a new pod isn't waiting on a network pull.

Build vs. buy, honestly

A hosted API is the correct engineering answer more often than infra people like to admit. You get no capacity planning, no GPU on-call, no cold starts, and someone else's continuous batching. At low or spiky volume, a self-hosted GPU sits idle burning money while your hosted bill tracks actual usage. Serving is a real specialty; "we'll just run vLLM" is a team commitment, not a weekend.

What genuinely justifies self-hosting:

  • Data residency or isolation that no vendor contract satisfies — regulated data, air-gapped deployment, an outright prohibition on third-party processing.
  • Sustained high volume. The economics invert when utilization is high and steady. Reserved GPUs win on cost only when they're busy; average utilization, not peak, is the number that decides this.
  • Custom weights. Fine-tuned, distilled, or otherwise modified models nobody hosts for you.
  • A latency floor you can't otherwise reach. Co-locating the model with your data, skipping a network hop, or guaranteeing capacity that isn't shared with other tenants.
  • Control over the serving stack — quantization scheme, sampling behaviour, deterministic versions that don't change under you.

What does not justify it: a vague preference for control, an assumption that self-hosting is cheaper without measuring utilization, or a proof of concept whose traffic is a rounding error.

A reasonable path: start hosted, instrument tokens in and out per request from day one, and revisit when sustained volume makes the arithmetic obvious. Keeping an OpenAI-compatible interface at your boundary makes that switch a config change rather than a rewrite — which is most of the argument for that interface shape in the first place.

Checklist

  • Know whether your traffic is prefill-dominated or decode-dominated; measure both token rates separately
  • Report TTFT, ITL, and end-to-end from a client, at concurrency 1 and at target concurrency
  • Size KV cache with num_kv_heads, not attention heads, and treat max_seq_len × concurrency as one shared budget
  • Use a serving engine with paged attention and continuous batching before optimizing anything else
  • Turn on prefix caching and keep shared prefixes byte-identical and leading
  • Use chunked prefill if long prompts are spiking ITL for other users
  • Quantize for memory first, speed second — and re-evaluate quality on your own traffic
  • Add speculative decoding only where batch sizes are genuinely low
  • Set explicit queue depth and admission limits; shed load rather than queueing unbounded
  • Scale on queue depth, keep a warm replica, and bake weights into the image
  • Reach for parallelism only when it won't fit; tensor within a node, pipeline across nodes

None of this is exotic — it's the same discipline as any other capacity-constrained service: know your bottleneck, measure what users feel, stop paying for headroom you never use. The twist is that with LLMs the bottleneck moves between two phases mid-request, and the cache, not the model, is what fills the card.

Start with continuous batching and paged attention. Measure. Then argue about quantization.