clock frequency
**Clock frequency** measured in **GHz determines the rate at which processors execute operations** — higher clock speeds mean more instructions per second, though modern AI workloads depend more on parallel throughput (FLOPS) and memory bandwidth than raw frequency.
**What Is Clock Frequency?**
- **Definition**: Number of clock cycles per second, measured in Hz/GHz.
- **Mechanism**: Each cycle, the processor advances through instruction stages.
- **Range**: Modern CPUs: 2-5+ GHz; GPUs: 1-2.5 GHz.
- **Relation**: Higher frequency generally equals faster single-thread performance.
**Why Frequency Matters**
- **Execution Speed**: More cycles = more operations per second.
- **Latency**: Faster clocks reduce time per operation.
- **Benchmark**: Common (if misleading) comparison metric.
- **Power**: Frequency directly impacts power consumption.
**Frequency vs. Performance**
**CPU Single-Thread**:
```
CPU | Base | Boost | Single-Thread Score
-----------------|----------|----------|--------------------
AMD 7950X | 4.5 GHz | 5.7 GHz | 2,100
Intel 14900K | 3.2 GHz | 6.0 GHz | 2,300
Apple M3 Max | 4.1 GHz | 4.1 GHz | 2,200
AMD 9950X | 4.3 GHz | 5.7 GHz | 2,300
```
**GPU Clocks**:
```
GPU | Base | Boost | Note
-----------------|----------|----------|-------------------
NVIDIA H100 | 1.1 GHz | 1.8 GHz | Lower than gaming
NVIDIA RTX 4090 | 2.2 GHz | 2.5 GHz | High consumer clock
AMD MI300X | 1.7 GHz | 2.1 GHz | Chiplet design
AMD RX 7900 XTX | 1.9 GHz | 2.5 GHz | High consumer clock
```
**Why GPU Clocks Are Lower**:
```
AI chips optimize for:
- Throughput (FLOPS) over latency
- Power efficiency
- Thermal sustainability
- Memory bandwidth
Gaming chips optimize for:
- Peak performance
- High clocks
- Short burst workloads
```
**FLOPS vs. Frequency**
**What Matters for AI**:
```
FLOPS = Clock × Cores × Operations/Clock
Example H100:
1.8 GHz × 16,896 SMs × 2 (FMA) × 128 (tensor cores) ≈ 1,979 TFLOPS (FP16)
Higher clocks help, but:
- Core count matters more
- Tensor cores multiply throughput
- Memory bandwidth is often the bottleneck
- Parallelism > frequency for AI
```
**Performance Formula**:
```
Single-thread: Frequency-sensitive
Parallel work: Core count × frequency
Memory-bound: Bandwidth-limited
AI inference: Memory bandwidth limited
AI training: Compute + bandwidth
```
**Frequency and Power**
**Power Relationship**:
```
Power ∝ Voltage² × Frequency
Higher frequency requires:
- Higher voltage
- More power
- More cooling
- Lower efficiency
Example:
5 GHz at 1.35V: 150W
4 GHz at 1.1V: 80W (47% less power)
```
**Efficiency Sweet Spot**:
```
Frequency | Power | Perf/Watt
-------------|--------|----------
100% (max) | 100% | 1.0
90% | 75% | 1.2
80% | 60% | 1.33
70% | 45% | 1.56
Often better to run lower frequency for efficiency
```
**Overclocking & Underclocking**
**For AI Workloads**:
```
Strategy | When to Use
----------------|----------------------------------
Default | Most production workloads
Overclock | Maximum performance (short runs)
Underclock | Efficiency, thermals, reliability
Power limit | Maintain perf while saving power
```
**GPU Power Limiting**:
```bash
# NVIDIA GPU power limit
nvidia-smi -pl 300 # Set to 300W (from 450W)
# Result: ~95% performance at 67% power
```
**Frequency Scaling**
**Dynamic Frequency**:
```
State | Frequency | When
----------------|--------------|-------------------
Idle | 300-500 MHz | No load
Base | 2-4 GHz | Sustained workload
Boost | 4-6 GHz | Thermal headroom
Thermal throttle|
Go deeper with CFSGPT
Get AI-powered deep-dives, save terms, and run advanced simulations — free account.
Create Free Account