GPU Streaming Multiprocessor (SM) Architecture is the fundamental execution unit of GPU chips, containing dozens of CUDA cores, tensor cores, warp schedulers, and hierarchical cache/memory subsystems orchestrated to achieve massive thread parallelism and memory bandwidth.
CUDA Core and Tensor Core Organization
- CUDA Cores: Scalar processing elements executing FP32 (single-precision) or integer operations. Typical SM: 32-128 CUDA cores. Each core contains FP unit, integer ALU, and special function unit (SFU).
- Tensor Cores: Specialized units performing matrix multiplication (4×4 or 8×8 matrix ops in few cycles). Recent GPUs (Volta+) dedicate substantial area to tensor cores (10-20 cores per SM).
- Special Function Units (SFU): Execute transcendental functions (sin, cos, reciprocal), integer operations. Typically 1 SFU per warp (32 threads) limiting throughput for special functions.
Warp Scheduling Hardware
- Warp Concept: Group of 32 threads executing in lockstep (SIMD). Modern GPUs issue 2-4 warps per cycle, each to different execution units.
- Warp Scheduler: Selects ready warps (no stalls) for execution from resident warps (typically 32-64 per SM). Scheduling policies: round-robin, priority-based, or two-level hierarchical.
- Ready Warp Identification: Tracks register availability, operand readiness, instruction fetch completion. Warp marked "stalled" when waiting for memory, synchronization, or resources.
- Dual-Issue Architecture: Modern designs issue two independent instructions from same warp or different warps. Enables pipelining and hiding latencies.
Register File Banking and Architecture
- Register File Size: 64-256 KB per SM (Ampere: 256 KB). Distributed as 32 banks, one read port per bank per cycle.
- Bank Conflict: Simultaneous accesses to same register bank by different threads. Causes serialization (pipeline stall) limiting throughput.
- Banking Layout: Registers allocated sequentially to threads. Thread i's registers in bank (i mod 32). Stride-1 accesses have no conflicts; stride-32 accesses fully serialize.
- Register Optimization: Compiler allocates registers to minimize bank conflicts. Unroll loops to increase register pressure but improve ILP. Register spilling to local memory expensive (~10x slower).
L1 Cache and Shared Memory Integration
- L1 Cache: 32-64 KB per SM. Caches all memory accesses (if enabled). Separate banks from shared memory in Ampere (flexible partitioning).
- Shared Memory: 48-96 KB fast on-chip memory, explicitly managed by programmer. Bank-conflict free access with properly aligned patterns (sequential access best).
- Write-Through Behavior: L1 write-through to L2 (no write-back buffering in early GPU architectures). Recent designs: write-back option for reduced memory traffic.
Load-Store Unit and Memory Subsystem
- Load-Store Capability: SM can issue multiple load/store instructions per cycle. Coalesced accesses (consecutive threads accessing consecutive memory addresses) merge into single bus transaction.
- Coalescing Efficiency: 32 consecutive loads (4-byte words) coalesce into one 128-byte transaction. Scattered patterns waste bandwidth.
- Memory Latency Hiding: 100-500 cycle memory latency hidden by scheduling other ready warps. Occupancy (resident warp count) determines latency hiding capability.
Occupancy and Latency Hiding
- Occupancy Metric: Percentage of maximum resident warps actually resident. Higher occupancy better hides memory latency (more warps available to schedule while others wait).
- Limiting Factors: Register pressure, shared memory allocation per thread, block size constraints determine max occupancy (typically 50-100%).
- Ampere/Hopper Evolution: Larger register files (256 KB), flexible shared memory partitioning, tensor float 32 (TF32) precision enable higher occupancy while maintaining performance.
gpu warp architecture designsm streaming multiprocessorcuda core executionregister file gpuwarp scheduler hardware design
Related Topics
Explore 500+ Semiconductor & AI Topics
From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.