Transformer Inference
# Transformer Inference and LLM Serving Architecture
Transformer inference is the production execution of a trained transformer to turn an input token sequence into predictions or generated tokens. For a large language model, serving is not one uniform computation. It has two phases with different hardware behavior: prefill processes the prompt in parallel and is usually compute-bound, while autoregressive decode emits one token per sequence per step and is usually constrained by memory bandwidth. A useful serving design treats latency, throughput, KV-cache capacity, numerical precision, scheduling, parallelism, and output quality as one system.
The user-visible path is:
1. tokenize and admit the request;
2. run prompt prefill and create the initial key-value cache;
3. repeatedly retrieve cached keys and values, compute the next-token logits, sample a token, and append its new cache state;
4. stream decoded text while measuring time to first token, inter-token latency, throughput, and tail latency.
## The two-phase workload
### Prefill: parallel prompt processing
During prefill, the model consumes all prompt tokens and produces the first set of attention keys and values. Matrix multiplications operate across many tokens at once, exposing enough parallel work to fill tensor cores. The arithmetic intensity, measured as operations performed per byte moved from memory, is much higher than during single-token decode. Large matrix-matrix multiplications can therefore approach the accelerator's compute ceiling when kernels, shapes, and tensor-parallel communication are efficient.
Prefill determines much of time to first token (TTFT). For a request admitted at time $t_0$ and a first output token emitted at $t_1$,
Prompt length changes the amount of prefill work. Dense self-attention has an $O(S^2)$ score matrix for sequence length $S$, while the projection and feed-forward layers scale approximately linearly with $S$ for a fixed model. FlashAttention does not change the exact attention result or its mathematical operation count; it tiles the computation so intermediate score and probability blocks stay in on-chip SRAM instead of being materialized repeatedly in HBM. That reduction in memory traffic can make long-prompt prefill substantially faster.
TTFT is not only model execution time. Admission queues, prefix-cache lookup, tokenization, distributed collectives, and cold model loading may dominate. Report queue and execution time separately at median, p95, and p99.
### Decode: sequential token generation
After prefill, autoregressive decode generates one new token per active sequence at each step. The new token becomes input to the next step, creating a true dependency chain. A single request exposes little matrix dimension in the token axis, so the accelerator must repeatedly read model weights and the growing KV cache to perform comparatively little arithmetic.
The roofline model expresses attainable performance as the smaller of the compute ceiling and the bandwidth ceiling:
where $P_{peak}$ is peak arithmetic throughput, $B_{HBM}$ is sustainable HBM bandwidth, and $I$ is arithmetic intensity in operations per byte. Prefill often lies closer to the compute roof. Low-batch decode has low $I$ and lies on the bandwidth slope. Adding more nominal FLOPS does not help when the memory channels are already saturated.
Inter-token latency (ITL) measures the delay between streamed output tokens. For output timestamps $t_i$ and $t_{i+1}$,
End-to-end latency for $N$ generated tokens is approximately
but averages hide pauses caused by batch reshaping, collective communication, preemption, or memory pressure. Interactive applications need an ITL distribution and stall rate, not only aggregate tokens per second.
## KV-cache memory hierarchy
Self-attention needs the keys and values from all prior tokens. Recomputing them at every step would make generation increasingly expensive, so serving systems store them in a KV cache. For batch size $B$, cached sequence length $S$, transformer layers $L$, key-value heads $H_{kv}$, per-head dimension $D_h$, and element width $b$ bytes, an approximate footprint is
The factor of two represents keys and values. Some descriptions use $d_{model}$ instead of $H_{kv}D_h$; that is correct for ordinary multi-head attention, but grouped-query attention and multi-query attention use fewer KV heads and therefore consume less cache memory. Allocator metadata, alignment, block tables, temporary workspaces, and fragmentation add overhead beyond the equation.
For example, 80 layers, 8 KV heads, head dimension 128, BF16 entries, and 8,192 cached tokens require roughly 2.5 GiB per sequence before overhead. Capacity determines concurrency: after weights and workspaces, safe active sequences are bounded by $\lfloor M_{budget}/M_{KV,request}\rfloor$, with headroom for variable outputs. SRAM holds active tiles, HBM holds weights and live KV blocks, and host memory is a slower swap tier. Admission based only on request count can exhaust memory when long-context requests arrive together.
### PagedAttention and block allocation
Naive serving reserves a contiguous worst-case KV region for each request. Actual output lengths vary, so reservations waste capacity and external fragmentation makes free memory unusable. PagedAttention divides the cache into fixed-size physical blocks and maps each sequence's logical token positions through a block table. A sequence can grow without one contiguous allocation, and completed sequences return blocks to a common pool.
Paging improves utilization; it does not make cache traffic free. Block size trades metadata overhead against internal fragmentation. Prefix sharing can reuse immutable system-prompt blocks, with reference counting and copy-on-write when sequences diverge. Eviction should consider priority and recomputation cost.
KV-cache quantization to FP8 or lower precision reduces capacity and bandwidth demand, but must be qualified for the model, attention distribution, context length, and output-quality target. Keys and values can have different sensitivity. A production report names cache dtype separately from weight dtype and activation dtype.
## Scheduling and continuous batching
Static batching waits for a group of requests, pads them to compatible shapes, and runs the group to completion. It is simple but wastes slots when sequences finish at different times. Continuous batching, used by serving engines such as vLLM and Text Generation Inference (TGI), updates the active batch at token boundaries: completed requests leave and queued requests enter without waiting for the longest sequence.
Larger decode batches reuse each weight load across more sequences, increasing aggregate throughput. The tradeoff is queueing and per-user ITL. Useful controls include:
- maximum batched tokens rather than only maximum sequences;
- separate limits for prefill tokens and decode sequences;
- chunked prefill so a long prompt does not block every decoder;
- age or deadline terms that prevent starvation;
- priority classes, memory-aware admission, and bounded preemption.
Prefill and decode contend for the same tensor cores, HBM, and interconnect. Combining them can improve utilization but add ITL jitter. Separate pools give tighter control but must transfer KV state and balance two fleets.
## Parallelism across accelerators
A model that does not fit, or does not meet latency on one accelerator, must be partitioned. Each parallelism method moves a different object and creates a different bottleneck.
### Tensor parallelism
Tensor parallelism (TP) splits matrices across devices. Each rank stores weight shards, computes partial results, and joins all-reduce or all-gather collectives. TP reduces per-device capacity and compute, but communicates within nearly every layer. More ranks are not automatically faster when synchronization dominates.
### Pipeline parallelism
Pipeline parallelism (PP) assigns consecutive layer groups to devices. It reduces per-device storage and can span weaker links, but each request traverses every stage serially. Microbatches fill the pipeline; unequal stage times create bubbles. Keep chatty TP collectives inside the fastest fabric, use PP across slower boundaries only when necessary, and measure collective tail latency under load.
## Precision and quantization
Inference precision changes weight capacity, memory traffic, arithmetic throughput, and output quality.
| Format | Typical use | Main benefit | Qualification concern |
|---|---|---|---|
| BF16/FP16 | baseline weights and activations | broad kernel support | weight bandwidth and capacity |
| FP8 | weights/activations on supported accelerators | higher tensor throughput, fewer bytes | scaling strategy and outliers |
| INT8 | weight or activation quantization | mature compression path | calibration coverage |
| INT4 | commonly weight-only | major capacity/bandwidth reduction | dequantization cost and quality loss |
Weight-only INT4 can roughly quarter weight bytes relative to BF16, allowing fewer devices and less decode traffic. Speedup depends on fused dequantization kernels, group size, shapes, and the next bottleneck. Qualify task quality, safety, long-context stability, and latency; perplexity alone is insufficient.
## Speculative decoding
Autoregressive dependency limits how many target-model tokens can be produced per serial step. Speculative decoding uses a cheaper draft model or auxiliary heads to propose several tokens, then verifies those proposals with the target model in parallel. Accepted prefixes advance the stream by multiple tokens while preserving the target distribution when the algorithm is implemented correctly.
Speedup depends on acceptance rate, draft cost, verification efficiency, and batch shape. A weak draft, distribution shift, or well-batched target can erase the benefit. Report accepted tokens per verification step and end-to-end latency, not just proposed tokens.
## Serving metrics and capacity planning
No single throughput number describes an inference service. Measure at least:
- TTFT and queue time at p50, p95, and p99;
- ITL and stall frequency for streamed generation;
- request throughput and generated tokens per second;
- prefill tokens per second separately from decode tokens per second;
- HBM used by weights, KV cache, workspaces, and fragmentation;
- batch size, cache occupancy, eviction rate, and HBM/interconnect utilization;
- cancellations, output lengths, deadline misses, quality, energy, and cost.
Capacity tests should replay representative prompt lengths, output lengths, arrivals, priorities, and cancellations. Fixed-length prompts hide fragmentation and tail effects. Report warm and cold behavior separately.
## Common failure modes
Optimizing average tokens per second while users wait. Large batches can improve fleet throughput and worsen TTFT or ITL. Enforce explicit latency budgets and priority-aware admission.
Using model size as the whole memory budget. KV cache, graph captures, communication buffers, allocator reserve, and runtime workspaces can determine whether the deployment fits.
Adding devices without modeling communication. TP collectives or expert all-to-all can turn an accelerator upgrade into an interconnect bottleneck.
Claiming quantization speedup from bytes alone. Unsupported kernels can add conversions and run slower. Measure the deployed path.
## Summary
Transformer inference is a hardware-software scheduling problem built around two regimes. Prefill offers parallel, compute-intensive work and drives TTFT. Decode is serial across tokens, repeatedly streams weights and KV state, and commonly drives ITL through HBM bandwidth. The KV cache converts recomputation into memory consumption; PagedAttention and prefix sharing improve how that memory is allocated. Continuous batching trades per-request latency against weight reuse. Tensor and pipeline parallelism trade device capacity against communication. FP8 and INT4 reduce bytes only when qualified kernels and acceptable model quality make the reduction useful. Speculative decoding attacks serial latency by verifying multiple proposed tokens at once.
The correct production configuration is not the one with the largest benchmark tokens-per-second number. It is the configuration that meets TTFT and ITL objectives at the required concurrency, fits weights and cache with recovery headroom, preserves output quality, and remains stable under the real distribution of prompts, outputs, cancellations, and failures.