Home Knowledge Base KV Cache and Prefix Caching

KV Cache and Prefix Caching

What is KV Cache? During autoregressive generation, the model computes key (K) and value (V) tensors for attention. Caching these avoids recomputation on each new token.

How KV Cache Works

Without Cache Every token generation recomputes attention for the entire sequence:

Token 1: Compute K,V for position 0
Token 2: Compute K,V for positions 0,1 (recompute!)
Token 3: Compute K,V for positions 0,1,2 (recompute!)

Quadratic complexity.

With Cache Store K,V from previous steps:

Token 1: Compute K,V for position 0, cache it
Token 2: Retrieve cached K,V, compute only position 1, append to cache
Token 3: Retrieve cached K,V, compute only position 2, append to cache

Linear complexity for generation.

KV Cache Size

Cache size = 2 × num_layers × seq_len × hidden_dim × num_kv_heads × bytes_per_param

Example for Llama-2 7B (BF16, 4K context):

Prefix Caching

The Problem Different requests often share common prefixes (system prompts):

Request 1: [System prompt] + [User query A]
Request 2: [System prompt] + [User query B]

Without caching: Recompute system prompt KV for every request.

With Prefix Caching Compute and cache system prompt KV once, reuse for all requests:

Prefix cache: System prompt KV (compute once)
Request 1: Reuse prefix + compute query A KV
Request 2: Reuse prefix + compute query B KV

PagedAttention (vLLM) Instead of contiguous memory for KV cache:

Physical blocks: [Block 0][Block 1][Block 2][Block 3]
Request 1 KV:    [  0   ][   1  ]
Request 2 KV:    [  0   ][   2  ] (shares prefix block 0)

Using Prefix Caching

vLLM

# Enable automatic prefix caching
python -m vllm.entrypoints.openai.api_server
    --model meta-llama/Llama-2-7b-chat-hf
    --enable-prefix-caching

Benefits

MetricWithout Prefix CacheWith Prefix Cache
TTFT~500ms~50ms (for cached prefix)
ThroughputBaseline2-3x higher
MemoryPer-requestShared

Prefix caching is especially valuable for chat applications with consistent system prompts.

kv cacheprefix cachingcache

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.