Home Knowledge Base Latency

Latency in the context of AI and LLM deployment refers to the time delay between sending a request to a model and receiving the beginning of its response. It is one of the most critical performance metrics for any real-time AI application.

Components of LLM Latency

Typical Latency Targets

Optimization Techniques

latencydeployment

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.