Register File in GPU architecture is a high-speed memory bank providing each thread with dedicated registers for storing operands and intermediate results.
What Is a Register File?
- Location: Inside each Streaming Multiprocessor (SM)
- Speed: Single-cycle access (fastest memory in GPU hierarchy)
- Capacity: 64KB-256KB per SM (varies by GPU generation)
- Allocation: Dynamically partitioned among threads
Why Register Files Matter
Register files enable thousands of concurrent threads by providing each thread private, zero-latency storage. Register pressure limits occupancy.
<svg viewBox="0 0 485 245" xmlns="http://www.w3.org/2000/svg" style="max-width:100%;height:auto" role="img"><rect x="0" y="0" width="485" height="245" rx="12" fill="#0d1117"/><g font-family="ui-monospace,SFMono-Regular,Menlo,Consolas,"Liberation Mono",monospace" font-size="14"><text xml:space="preserve" x="20" y="31.7"><tspan fill="#c9d1d9">GPU Memory Hierarchy (NVIDIA):</tspan></text><text xml:space="preserve" x="20" y="50.7"><tspan fill="#6e7681">┌─────────────────────────────────────────┐</tspan></text><text xml:space="preserve" x="20" y="69.7"><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> Register File (per thread) </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> </tspan><tspan fill="#6e7681">←</tspan><tspan fill="#c9d1d9"> Fastest</tspan></text><text xml:space="preserve" x="20" y="88.7"><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> 255 registers × 4 bytes = 1KB per thread</tspan><tspan fill="#6e7681">│</tspan></text><text xml:space="preserve" x="20" y="107.7"><tspan fill="#6e7681">├─────────────────────────────────────────┤</tspan></text><text xml:space="preserve" x="20" y="126.7"><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> Shared Memory (per block) - 48-163KB </tspan><tspan fill="#6e7681">│</tspan></text><text xml:space="preserve" x="20" y="145.7"><tspan fill="#6e7681">├─────────────────────────────────────────┤</tspan></text><text xml:space="preserve" x="20" y="164.7"><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> L1/L2 Cache </tspan><tspan fill="#6e7681">│</tspan></text><text xml:space="preserve" x="20" y="183.7"><tspan fill="#6e7681">├─────────────────────────────────────────┤</tspan></text><text xml:space="preserve" x="20" y="202.7"><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> Global Memory (GDDR/HBM) - 8-80GB </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> </tspan><tspan fill="#6e7681">←</tspan><tspan fill="#c9d1d9"> Slowest</tspan></text><text xml:space="preserve" x="20" y="221.7"><tspan fill="#6e7681">└─────────────────────────────────────────┘</tspan></text></g></svg>
Register Pressure Trade-off:
| Registers/Thread | Threads/SM | Occupancy |
|---|---|---|
| 32 | 2048 | 100% |
| 64 | 1024 | 50% |
| 128 | 512 | 25% |
Fewer registers = more threads, but may cause spills to slow memory.
register filegpu registersthread storage
Explore 500+ Semiconductor & AI Topics
From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.