Home Knowledge Base System design for LLM applications

System design for LLM applications involves architecting scalable, reliable infrastructure to serve AI capabilities to users — addressing unique challenges like variable latency, high memory requirements, and non-deterministic outputs while applying traditional system design principles for load balancing, caching, and fault tolerance.

What Is LLM System Design?

Why LLM System Design Differs

High-Level Architecture

<svg viewBox="0 0 569 625" xmlns="http://www.w3.org/2000/svg" style="max-width:100%;height:auto" role="img"><rect x="0" y="0" width="569" height="625" rx="12" fill="#0d1117"/><g font-family="ui-monospace,SFMono-Regular,Menlo,Consolas,&quot;Liberation Mono&quot;,monospace" font-size="14"><text xml:space="preserve" x="20" y="31.7"><tspan fill="#6e7681">┌─────────────────────────────────────────────────────────────┐</tspan></text><text xml:space="preserve" x="20" y="50.7"><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9">                        Clients                              </tspan><tspan fill="#6e7681">│</tspan></text><text xml:space="preserve" x="20" y="69.7"><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9">              (Web, Mobile, API consumers)                   </tspan><tspan fill="#6e7681">│</tspan></text><text xml:space="preserve" x="20" y="88.7"><tspan fill="#6e7681">└─────────────────────────────────────────────────────────────┘</tspan></text><text xml:space="preserve" x="20" y="107.7"><tspan fill="#c9d1d9">                            </tspan><tspan fill="#6e7681">│</tspan></text><text xml:space="preserve" x="20" y="126.7"><tspan fill="#c9d1d9">                            </tspan><tspan fill="#6e7681">▼</tspan></text><text xml:space="preserve" x="20" y="145.7"><tspan fill="#6e7681">┌─────────────────────────────────────────────────────────────┐</tspan></text><text xml:space="preserve" x="20" y="164.7"><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9">                    Load Balancer                            </tspan><tspan fill="#6e7681">│</tspan></text><text xml:space="preserve" x="20" y="183.7"><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9">              (nginx, AWS ALB, CloudFlare)                   </tspan><tspan fill="#6e7681">│</tspan></text><text xml:space="preserve" x="20" y="202.7"><tspan fill="#6e7681">└─────────────────────────────────────────────────────────────┘</tspan></text><text xml:space="preserve" x="20" y="221.7"><tspan fill="#c9d1d9">                            </tspan><tspan fill="#6e7681">│</tspan></text><text xml:space="preserve" x="20" y="240.7"><tspan fill="#c9d1d9">                            </tspan><tspan fill="#6e7681">▼</tspan></text><text xml:space="preserve" x="20" y="259.7"><tspan fill="#6e7681">┌─────────────────────────────────────────────────────────────┐</tspan></text><text xml:space="preserve" x="20" y="278.7"><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9">                    API Gateway                              </tspan><tspan fill="#6e7681">│</tspan></text><text xml:space="preserve" x="20" y="297.7"><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9">    - Rate limiting                                          </tspan><tspan fill="#6e7681">│</tspan></text><text xml:space="preserve" x="20" y="316.7"><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9">    - Authentication                                         </tspan><tspan fill="#6e7681">│</tspan></text><text xml:space="preserve" x="20" y="335.7"><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9">    - Request routing                                        </tspan><tspan fill="#6e7681">│</tspan></text><text xml:space="preserve" x="20" y="354.7"><tspan fill="#6e7681">└─────────────────────────────────────────────────────────────┘</tspan></text><text xml:space="preserve" x="20" y="373.7"><tspan fill="#c9d1d9">                            </tspan><tspan fill="#6e7681">│</tspan></text><text xml:space="preserve" x="20" y="392.7"><tspan fill="#c9d1d9">              </tspan><tspan fill="#6e7681">┌─────────────┼─────────────┐</tspan></text><text xml:space="preserve" x="20" y="411.7"><tspan fill="#c9d1d9">              </tspan><tspan fill="#6e7681">▼</tspan><tspan fill="#c9d1d9">             </tspan><tspan fill="#6e7681">▼</tspan><tspan fill="#c9d1d9">             </tspan><tspan fill="#6e7681">▼</tspan></text><text xml:space="preserve" x="20" y="430.7"><tspan fill="#6e7681">┌─────────────────┐</tspan><tspan fill="#c9d1d9"> </tspan><tspan fill="#6e7681">┌─────────────┐</tspan><tspan fill="#c9d1d9"> </tspan><tspan fill="#6e7681">┌─────────────────┐</tspan></text><text xml:space="preserve" x="20" y="449.7"><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9">  Cache Layer    </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> RAG Service </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9">  LLM Service    </tspan><tspan fill="#6e7681">│</tspan></text><text xml:space="preserve" x="20" y="468.7"><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9">  (Redis)        </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> (Retrieval) </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9">  (vLLM/TGI)     </tspan><tspan fill="#6e7681">│</tspan></text><text xml:space="preserve" x="20" y="487.7"><tspan fill="#6e7681">└─────────────────┘</tspan><tspan fill="#c9d1d9"> </tspan><tspan fill="#6e7681">└─────────────┘</tspan><tspan fill="#c9d1d9"> </tspan><tspan fill="#6e7681">└─────────────────┘</tspan></text><text xml:space="preserve" x="20" y="506.7"><tspan fill="#c9d1d9">                            </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9">             </tspan><tspan fill="#6e7681">│</tspan></text><text xml:space="preserve" x="20" y="525.7"><tspan fill="#c9d1d9">                            </tspan><tspan fill="#6e7681">▼</tspan><tspan fill="#c9d1d9">             </tspan><tspan fill="#6e7681">▼</tspan></text><text xml:space="preserve" x="20" y="544.7"><tspan fill="#c9d1d9">                    </tspan><tspan fill="#6e7681">┌─────────────┐</tspan><tspan fill="#c9d1d9"> </tspan><tspan fill="#6e7681">┌─────────────────┐</tspan></text><text xml:space="preserve" x="20" y="563.7"><tspan fill="#c9d1d9">                    </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> Vector DB   </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9">  GPU Cluster    </tspan><tspan fill="#6e7681">│</tspan></text><text xml:space="preserve" x="20" y="582.7"><tspan fill="#c9d1d9">                    </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> (Pinecone)  </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9">  (H100s)        </tspan><tspan fill="#6e7681">│</tspan></text><text xml:space="preserve" x="20" y="601.7"><tspan fill="#c9d1d9">                    </tspan><tspan fill="#6e7681">└─────────────┘</tspan><tspan fill="#c9d1d9"> </tspan><tspan fill="#6e7681">└─────────────────┘</tspan></text></g></svg>

Key Components

API Layer:

from fastapi import FastAPI, HTTPException
from fastapi.middleware.cors import CORSMiddleware
import asyncio

app = FastAPI()

@app.post("/v1/chat/completions")
async def chat_completion(request: ChatRequest):
    # Validate
    if len(request.messages) == 0:
        raise HTTPException(400, "Messages required")
    
    # Check cache
    cache_key = hash_request(request)
    if cached := await cache.get(cache_key):
        return cached
    
    # Route to appropriate model
    model_endpoint = get_model_endpoint(request.model)
    
    # Generate (with timeout)
    try:
        response = await asyncio.wait_for(
            generate(model_endpoint, request),
            timeout=60.0
        )
    except asyncio.TimeoutError:
        raise HTTPException(504, "Generation timeout")
    
    # Cache and return
    await cache.set(cache_key, response, ttl=3600)
    return response

Scaling Strategies

Horizontal Scaling:

Load Pattern          | Strategy
----------------------|----------------------------------
Bursty traffic        | Auto-scaling GPU instances
Predictable peaks     | Scheduled scaling
Global users          | Multi-region deployment
Cost optimization     | Spot instances + fallback

Caching Layers:

Layer              | Cache What           | TTL
-------------------|----------------------|----------
Response cache     | Full responses       | 1-24 hours
Embedding cache    | Vector embeddings    | Days
KV cache           | Attention states     | Session
Prefix cache       | System prompts       | Hours

Multi-Model Routing

def route_request(request):
    """Route to appropriate model based on complexity."""
    
    # Simple queries → small/fast model
    if is_simple_query(request):
        return "gpt-4o-mini"
    
    # Complex reasoning → large model
    if needs_complex_reasoning(request):
        return "gpt-4o"
    
    # Default
    return "gpt-4o-mini"

def is_simple_query(request):
    prompt = request.messages[-1].content
    return (
        len(prompt) < 100 and
        not any(word in prompt for word in 
                ["explain", "analyze", "compare"])
    )

Reliability Patterns

Circuit Breaker:

class CircuitBreaker:
    def __init__(self, failure_threshold=5, reset_timeout=60):
        self.failures = 0
        self.state = "closed"
        self.last_failure = None
    
    async def call(self, func, *args):
        if self.state == "open":
            if time.time() - self.last_failure > self.reset_timeout:
                self.state = "half-open"
            else:
                raise CircuitOpenError()
        
        try:
            result = await func(*args)
            self.failures = 0
            self.state = "closed"
            return result
        except Exception as e:
            self.failures += 1
            self.last_failure = time.time()
            if self.failures >= self.failure_threshold:
                self.state = "open"
            raise

Fallback Chain:

async def generate_with_fallback(request):
    providers = ["openai", "anthropic", "local"]
    
    for provider in providers:
        try:
            return await generate(provider, request)
        except Exception as e:
            logger.warning(f"{provider} failed: {e}")
            continue
    
    raise AllProvidersFailedError()

Monitoring & Observability

Metric              | What to Track
--------------------|--------------------------------
Latency (P50/P95)   | TTFT, total generation time
Throughput          | Requests/sec, tokens/sec
Error rate          | 4xx, 5xx, timeouts
GPU utilization     | Memory, compute usage
Cost                | Tokens per request, $/query

System design for LLM applications requires balancing performance, reliability, and cost — applying proven distributed systems patterns while adapting to the unique constraints of GPU-bound, high-latency inference workloads.

system designarchitecturescalingload balancingcachingreliabilityllm infrastructure

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.