AI Engineering / Inference / AI Infrastructure / LLMs / Performance
Inference Engineering: What Happens Between the Prompt and the Token
Prefill and decode, time to first token and time per token, the KV cache, continuous batching, prefix caching, quantization, and speculative decoding, plus what application teams calling a model API can control themselves.
On this page
From the outside, a language model API looks like any other request/response endpoint: send a prompt, receive text. Underneath, generation has a structure that determines latency, throughput, and cost in ways that ordinary web services do not prepare you for. Inference engineering is the work of serving models efficiently, and, for teams that call someone else's API, of shaping requests so the serving system can do its job.
This article covers the mechanics first, then what each side can control.
Two phases: prefill and decode
Generating a response happens in two distinct phases.
Prefill processes the entire prompt. Every prompt token can be processed in parallel, so this phase does a lot of computation in large, efficient batches of matrix multiplication. It tends to be compute-bound. Its output is the first generated token, plus the cached attention state for every prompt token.
Decode generates the rest of the output one token at a time. Each step depends on the previous token, so within one sequence the steps are strictly sequential. Each step does little computation per sequence but must read the model's weights, and the sequence's cached attention state, from GPU memory. At small batch sizes, decode is memory-bandwidth-bound: the hardware spends its time moving bytes, not multiplying them.
That asymmetry explains most of what follows. Long prompts mostly cost prefill time. Long outputs mostly cost decode steps. And batching many sequences together is how servers make decode efficient, because one read of the weights can serve every sequence in the batch.
The latency metrics that matter
A single "latency" number hides the structure. Measure these separately:
- Time to first token (TTFT): queueing plus prefill. This is what users perceive as responsiveness in a streaming interface.
- Time per output token (TPOT), also called inter-token latency: the decode step time. This determines how fast text appears once it starts.
- End-to-end latency for a response of n tokens is approximately TTFT + (n − 1) × TPOT.
- Throughput: total tokens generated per second across all requests. It trades against per-request latency, because bigger batches raise throughput and lengthen each step.
A useful framing for capacity is goodput: the throughput achieved while still meeting latency targets for TTFT and TPOT. A server can post a high raw throughput number by running huge batches that make every individual user wait.
The KV cache
During attention, every token's key and value vectors are needed by every later token. Recomputing them at each decode step would be wasteful, so servers keep them in GPU memory: the KV cache.
Its size per token follows directly from the model's shape:
KV bytes per token = 2 (K and V) × layers × KV heads × head dimension × bytes per valueTake a hypothetical model configuration with 32 layers, 8 key-value heads, a head dimension of 128, and 16-bit values:
2 × 32 × 8 × 128 × 2 bytes = 131,072 bytes = 128 KiB per token
4,096-token sequence → 512 MiB
32 such sequences at once → 16 GiBThe exact numbers depend on the model, but the lesson generalizes: KV cache memory, not compute, is often what limits how many requests a server can run concurrently, and it grows with both context length and concurrency. Architectural choices like grouped-query attention, where several query heads share fewer key-value heads, exist largely to shrink it.
Batching: static and continuous
The simplest batching groups requests, runs them together until all finish, and then starts the next group. Because outputs have different lengths, short sequences finish early and their slots sit idle until the longest one completes, while new requests wait in the queue.
Continuous batching, also called iteration-level scheduling and described in the Orca paper (OSDI 2022), makes scheduling decisions at every decode step. Finished sequences leave the batch immediately, and waiting sequences join.
Continuous batching makes the KV cache's memory layout matter. Sequences of unpredictable length come and go constantly, and reserving contiguous memory for each sequence's maximum length wastes most of it. vLLM's PagedAttention (SOSP 2023) borrows from virtual memory: the cache is allocated in fixed-size blocks, mapped per sequence, so memory is used close to what sequences actually need, and identical prefixes can share blocks.
Two further refinements are common in modern servers. Chunked prefill splits a long prompt's prefill into pieces interleaved with other sequences' decode steps, so one huge prompt does not stall every stream. Disaggregated serving runs prefill and decode on separate pools of hardware, because the two phases want different resources.
Prefix caching
Many requests share a long common beginning: the same system instructions, the same tool definitions, the same document being asked several questions. If the server keeps the KV cache for a prefix it has already processed, a later request with the same prefix can skip that part of prefill entirely.
Self-hosted servers implement this as automatic prefix caching, and several model API providers offer prompt caching, typically with reduced pricing and latency for cached input. The application-side implication is the same either way: put stable content first and variable content last. A timestamp or user ID near the top of the prompt makes every request's prefix unique and defeats the cache.
Making decode cheaper
Because decode is dominated by memory traffic, the main levers reduce bytes moved or the number of sequential steps.
- Quantization stores weights, and sometimes the KV cache, in fewer bits, for example 8-bit or 4-bit instead of 16-bit. Fewer bytes per step means faster decode and room for larger batches. The cost is some loss of accuracy, which varies by model, method, and task, and has to be measured on your own evaluation set rather than assumed.
- Speculative decoding uses a small, fast draft model to propose several tokens, which the large model then verifies in a single forward pass. With the acceptance procedure described by Leviathan, Kalman, and Matias (ICML 2023), the output distribution matches the large model's exactly. The speedup depends on how often the draft's guesses are accepted, so it is largest on predictable text.
- Smaller models for easier work. Routing simple requests, such as classification, extraction, and short rewrites, to a smaller model is often the largest cost and latency improvement available, and it is an application decision.
What application teams control
Most teams consume inference through an API and never touch a GPU. They still control much of the latency and cost their users experience.
- Prompt size. Every input token costs prefill time and money. Trim retrieved context, compact conversation history, and cut instructions the model does not need.
- Output length. Set a maximum output token count that fits the task, and ask for concise formats. Output tokens are paid for with decode steps.
- Prompt structure for caching. Stable instructions and tool definitions first, request-specific content last.
- Streaming. For interactive use, stream tokens. Perceived latency becomes TTFT instead of end-to-end latency.
- Timeouts that match generation. A fixed total timeout either cuts off long valid responses or waits far too long for stalled ones. Time out on TTFT and on gaps between tokens instead.
- Rate limits and overload. Treat
429and overload responses as backpressure: back off with jitter, respect anyretry-afterhint, and shed or queue low-priority work. - Model routing. Use the smallest model that passes your evaluation for each task, and keep the routing decision in code where it can be measured.
- Caching results. Identical requests with deterministic settings, such as classification of the same text, can be cached like any other function result.
A streaming client that measures the two latency components and enforces both timeouts is short:
async function streamCompletion(
request: CompletionRequest,
{ firstTokenTimeoutMs = 10_000, stallTimeoutMs = 5_000 } = {}
) {
const controller = new AbortController();
const startedAt = performance.now();
let lastTokenAt = startedAt;
let firstTokenAt: number | undefined;
let tokens = 0;
// One timer covers both phases: waiting for the first token, then gaps.
const watchdog = setInterval(() => {
const now = performance.now();
const limit = firstTokenAt === undefined ? firstTokenTimeoutMs : stallTimeoutMs;
if (now - lastTokenAt > limit) controller.abort(new Error('stream stalled'));
}, 250);
try {
for await (const chunk of client.stream(request, { signal: controller.signal })) {
const now = performance.now();
firstTokenAt ??= now;
lastTokenAt = now;
tokens += chunk.tokenCount;
yieldToUi(chunk.text);
}
} finally {
clearInterval(watchdog);
}
const ttftMs = (firstTokenAt ?? performance.now()) - startedAt;
const decodeMs = lastTokenAt - (firstTokenAt ?? lastTokenAt);
metrics.record({
ttftMs,
tpotMs: tokens > 1 ? decodeMs / (tokens - 1) : undefined,
outputTokens: tokens,
model: request.model,
});
}Recording TTFT and TPOT per model and per feature turns "the AI feature feels slow" into a diagnosable question. Long TTFT points at prompt size, queueing, or cache misses. Slow TPOT points at model choice or provider load. Large output counts point at prompts that invite verbosity.
Putting it together
Inference performance is not a single number to optimize. It is a set of tradeoffs between responsiveness, throughput, cost, and quality, decided at several layers: the model's architecture, the serving system's scheduling and memory management, and the application's prompts, routing, and timeouts. The serving layer can make every token cheaper. The application layer decides how many tokens are needed.
The surrounding software that decides what goes into each request, and how runs are recorded and evaluated, is covered in Harness Engineering.
Further reading
- Gyeong-In Yu et al., "Orca: A Distributed Serving System for Transformer-Based Generative Models" (OSDI 2022).
- Woosuk Kwon et al., "Efficient Memory Management for Large Language Model Serving with PagedAttention" (SOSP 2023).
- Yaniv Leviathan, Matan Kalman, and Yossi Matias, "Fast Inference from Transformers via Speculative Decoding" (ICML 2023).
- Joshua Ainslie et al., "GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints" (2023).