clock frequency

**Clock frequency** measured in **GHz determines the rate at which processors execute operations** — higher clock speeds mean more instructions per second, though modern AI workloads depend more on parallel throughput (FLOPS) and memory bandwidth than raw frequency. **What Is Clock Frequency?** - **Definition**: Number of clock cycles per second, measured in Hz/GHz. - **Mechanism**: Each cycle, the processor advances through instruction stages. - **Range**: Modern CPUs: 2-5+ GHz; GPUs: 1-2.5 GHz. - **Relation**: Higher frequency generally equals faster single-thread performance. **Why Frequency Matters** - **Execution Speed**: More cycles = more operations per second. - **Latency**: Faster clocks reduce time per operation. - **Benchmark**: Common (if misleading) comparison metric. - **Power**: Frequency directly impacts power consumption. **Frequency vs. Performance** **CPU Single-Thread**: ``` CPU | Base | Boost | Single-Thread Score -----------------|----------|----------|-------------------- AMD 7950X | 4.5 GHz | 5.7 GHz | 2,100 Intel 14900K | 3.2 GHz | 6.0 GHz | 2,300 Apple M3 Max | 4.1 GHz | 4.1 GHz | 2,200 AMD 9950X | 4.3 GHz | 5.7 GHz | 2,300 ``` **GPU Clocks**: ``` GPU | Base | Boost | Note -----------------|----------|----------|------------------- NVIDIA H100 | 1.1 GHz | 1.8 GHz | Lower than gaming NVIDIA RTX 4090 | 2.2 GHz | 2.5 GHz | High consumer clock AMD MI300X | 1.7 GHz | 2.1 GHz | Chiplet design AMD RX 7900 XTX | 1.9 GHz | 2.5 GHz | High consumer clock ``` **Why GPU Clocks Are Lower**: ``` AI chips optimize for: - Throughput (FLOPS) over latency - Power efficiency - Thermal sustainability - Memory bandwidth Gaming chips optimize for: - Peak performance - High clocks - Short burst workloads ``` **FLOPS vs. Frequency** **What Matters for AI**: ``` FLOPS = Clock × Cores × Operations/Clock Example H100: 1.8 GHz × 16,896 SMs × 2 (FMA) × 128 (tensor cores) ≈ 1,979 TFLOPS (FP16) Higher clocks help, but: - Core count matters more - Tensor cores multiply throughput - Memory bandwidth is often the bottleneck - Parallelism > frequency for AI ``` **Performance Formula**: ``` Single-thread: Frequency-sensitive Parallel work: Core count × frequency Memory-bound: Bandwidth-limited AI inference: Memory bandwidth limited AI training: Compute + bandwidth ``` **Frequency and Power** **Power Relationship**: ``` Power ∝ Voltage² × Frequency Higher frequency requires: - Higher voltage - More power - More cooling - Lower efficiency Example: 5 GHz at 1.35V: 150W 4 GHz at 1.1V: 80W (47% less power) ``` **Efficiency Sweet Spot**: ``` Frequency | Power | Perf/Watt -------------|--------|---------- 100% (max) | 100% | 1.0 90% | 75% | 1.2 80% | 60% | 1.33 70% | 45% | 1.56 Often better to run lower frequency for efficiency ``` **Overclocking & Underclocking** **For AI Workloads**: ``` Strategy | When to Use ----------------|---------------------------------- Default | Most production workloads Overclock | Maximum performance (short runs) Underclock | Efficiency, thermals, reliability Power limit | Maintain perf while saving power ``` **GPU Power Limiting**: ```bash # NVIDIA GPU power limit nvidia-smi -pl 300 # Set to 300W (from 450W) # Result: ~95% performance at 67% power ``` **Frequency Scaling** **Dynamic Frequency**: ``` State | Frequency | When ----------------|--------------|------------------- Idle | 300-500 MHz | No load Base | 2-4 GHz | Sustained workload Boost | 4-6 GHz | Thermal headroom Thermal throttle|

Go deeper with CFSGPT

Get AI-powered deep-dives, save terms, and run advanced simulations — free account.

Create Free Account