Streaming Multiprocessor (SM) is the fundamental compute building block in NVIDIA GPU architecture, containing CUDA cores, tensor cores, and shared resources for parallel execution.
What Is a Streaming Multiprocessor?
- Components: CUDA cores, tensor cores, LD/ST units, SFUs
- Resources: Registers, shared memory, L1 cache
- Scheduling: Multiple warps execute concurrently
- Scale: Consumer GPUs: 20-80 SMs; Data center: 100+ SMs
Why SM Architecture Matters
Understanding SM organization is essential for GPU programming optimization. Performance depends on efficiently utilizing SM resources.
<svg viewBox="0 0 401 321" xmlns="http://www.w3.org/2000/svg" style="max-width:100%;height:auto" role="img"><rect x="0" y="0" width="401" height="321" rx="12" fill="#0d1117"/><g font-family="ui-monospace,SFMono-Regular,Menlo,Consolas,"Liberation Mono",monospace" font-size="14"><text xml:space="preserve" x="20" y="31.7"><tspan fill="#c9d1d9">Streaming Multiprocessor (SM) Structure:</tspan></text><text xml:space="preserve" x="20" y="50.7"><tspan fill="#6e7681">┌─────────────────────────────────────────┐</tspan></text><text xml:space="preserve" x="20" y="69.7"><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> Instruction Cache </tspan><tspan fill="#6e7681">│</tspan></text><text xml:space="preserve" x="20" y="88.7"><tspan fill="#6e7681">├─────────────────────────────────────────┤</tspan></text><text xml:space="preserve" x="20" y="107.7"><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> Warp Scheduler (4 per SM) </tspan><tspan fill="#6e7681">│</tspan></text><text xml:space="preserve" x="20" y="126.7"><tspan fill="#6e7681">├─────────┬─────────┬─────────┬──────────┤</tspan></text><text xml:space="preserve" x="20" y="145.7"><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9">64 CUDA </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9">64 CUDA </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9">4 Tensor </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9">8 SFU </tspan><tspan fill="#6e7681">│</tspan></text><text xml:space="preserve" x="20" y="164.7"><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9">Cores </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9">Cores </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9">Cores </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9">Units </tspan><tspan fill="#6e7681">│</tspan></text><text xml:space="preserve" x="20" y="183.7"><tspan fill="#6e7681">├─────────┴─────────┴─────────┴──────────┤</tspan></text><text xml:space="preserve" x="20" y="202.7"><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> Load/Store Units (32) </tspan><tspan fill="#6e7681">│</tspan></text><text xml:space="preserve" x="20" y="221.7"><tspan fill="#6e7681">├─────────────────────────────────────────┤</tspan></text><text xml:space="preserve" x="20" y="240.7"><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> Register File (64K 32-bit registers) </tspan><tspan fill="#6e7681">│</tspan></text><text xml:space="preserve" x="20" y="259.7"><tspan fill="#6e7681">├─────────────────────────────────────────┤</tspan></text><text xml:space="preserve" x="20" y="278.7"><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> Shared Memory / L1 Cache (128KB) </tspan><tspan fill="#6e7681">│</tspan></text><text xml:space="preserve" x="20" y="297.7"><tspan fill="#6e7681">└─────────────────────────────────────────┘</tspan></text></g></svg>
SM Evolution (NVIDIA):
| Architecture | SMs (Max) | CUDA Cores/SM |
|---|---|---|
| Pascal | 60 | 64 |
| Volta | 84 | 64 |
| Ampere | 108 | 64 |
| Hopper | 132 | 128 |
streaming multiprocessorsmgpu architecture
Explore 500+ Semiconductor & AI Topics
From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.