Home Knowledge Base L2 cache

L2 cache is the shared on-chip cache level that serves all streaming multiprocessors before traffic reaches high-latency HBM - it acts as a global reuse and coherence layer for data exchanged across blocks and kernels on the same GPU.

What Is L2 cache?

Why L2 cache Matters

How It Is Used in Practice

L2 cache is the shared memory-traffic stabilizer for GPU-wide execution - strong L2 locality can significantly raise end-to-end kernel performance.

l2 cachel2hardware

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.