DRAM Ecc
# DRAM ECC: Hamming SECDED Architecture, Reliability Theory, and the Chipkill Application
ECC is the layer that stops pretending the physics underneath it is perfect. Every mechanism this series has traced — the storage capacitor's leak, the STI corner's parasitic path, the buried word line's GIDL trade-off, the sense amplifier's shrinking margin, row hammer's accelerated disturbance — occasionally still loses. ECC does not fix any of that. It adds pure mathematics on top of imperfect physics: extra redundant bits, computed and checked on every single access, that can catch and repair the errors every layer below it ultimately failed to prevent.
That 72-bit width is not an arbitrary round number — it is the smallest codeword the math allows. For $r$ check bits to uniquely identify an error among $k$ data bits (including the possibility of no error at all), the Hamming bound requires:
For $k = 64$, $r = 7$ check bits are the minimum that satisfies this ($2^7 = 128 \geq 64+7+1=72$), which gives plain single-error correction. Adding one more bit as an overall parity check over the whole codeword costs nothing in data width but raises the minimum Hamming distance from 3 to 4, upgrading the code from SEC to SECDED — single error correction, double error detection — which is why the real-world number is 8 check bits, not 7.
None of that changes what happens upstream — it only changes what happens to the result. Feed the exact same per-bit error rate this series has spent its entire length trying to reduce — leakage past the refresh deadline, an STI corner leak, a GIDL-accelerated access transistor, a disturbed row-hammer victim — into a word protected by SECDED, and the single-bit failures that used to corrupt data silently instead get caught and repaired automatically; only the much rarer case of two independent errors landing in the same 64-bit word at once gets through as a flagged, uncorrectable error instead of a silent one.
That residual gap — what happens when two errors land in the same word, or an entire chip dies at once — is exactly what plain SECDED was never built to survive. A standard ECC DIMM spreads one 72-bit codeword across 9 separate chips, so a normal x8 chip failure dumps 8 wrong bits into a single codeword at once — far more than SECDED's one-bit correction budget can absorb. Server memory solves this by changing how the codeword is built, not just how wide it is.
That reframing — redefining the unit of error rather than chasing the physics that causes it — is the real idea ECC contributes to this entire series. Every other keyword here was about delaying, detecting, or outracing a physical failure at the cell or circuit level. ECC doesn't compete with any of that; it sits above all of it, accepting that the capacitor will leak, the trench will occasionally let a corner leak, the word line will occasionally misfire, and simply making sure the system built on top never has to find out.
Read DRAM ECC through a *redundancy-is-cheaper-than-perfection* lens rather than a *"magic error fixer"* lens: the Hamming bound $2^r \geq k+r+1$ sets the minimum tax redundancy has to pay, SECDED's single-bit correction absorbs almost everything this series' physical mechanisms throw at it, and Chipkill is simply that same bargain renegotiated so an entire failed chip counts as one error instead of eight. Nothing below this layer got more reliable; the system above it simply stopped needing it to be.