DRAM Refresh Controller

# DRAM Refresh Controller: Row-Scheduling Architecture, the Refresh-Tax Theory, and Per-Bank Refresh Application

The refresh controller is the part of DRAM that never gets to rest. Every cell in the array is leaking charge continuously, and the controller's entire job is to read and rewrite every single row before any of them decays past the point the sense amplifier can still resolve it — forever, for as long as the chip is powered, interleaved with whatever real reads and writes the rest of the system is trying to do at the same time. It doesn't fix any of the leakage this project has traced through the storage capacitor, the isolation trench, or the buried word line; it just outruns all of it, continuously, on a schedule that gets tighter every generation.

Inside the Refresh Controller — Racing a Clock Against a Leak every row must be touched before the window closes, interleaved with real reads and writes row address counter cycles 0 … R−1 tREFI timer fires on a fixed interval on-die temp sensor drives the rate multiplier ×2 above 85°C REF command generator pending R/W commands command arbiter refresh can be postponed briefly to let a read finish, but it never loses the race indefinitely the arbiter can let a few real reads or writes go first, borrowing a little time from the refresh schedule — but every row still has to be serviced before the window closes, so delay here is always paid back later, not avoided

How often that timer has to fire is fixed by simple division. Every row in the array must be refreshed at least once within the retention window $t_{REFW}$ established back in the 1T1C cell's own decay equation, so if refresh commands are issued one at a time, each one covering its own slice of rows, the interval between them is just:

$$ t_{REFI} = \frac{t_{REFW}}{N} $$

where $N$ is the number of refresh commands needed to cover the whole array once. What actually costs bandwidth is not $t_{REFI}$ itself but $t_{RFC}$ — the time each individual refresh command occupies the array, unavailable for anything else — and the fraction of all time spent refreshing instead of serving real traffic is simply:

$$ \text{refresh overhead} = \frac{t_{RFC}}{t_{REFI}} $$
The Refresh Tax Climbs With Every Density Generation t_RFC grows with density; t_REFI barely does — the ratio can only go one way DRAM generation / density → refresh overhead (% of array time) → lower density higher density t_RFC lengthens every node — more rows per command, more array to cycle refresh overhead = t_RFC / t_REFI — and the same Arrhenius doubling from the 1T1C cell's own decay equation forces the controller to double its rate above 85°C, pushing the tax even higher exactly when it hurts most

That climbing tax is why "refresh all banks at once, then let the chip sit idle for the rest of the interval" stopped being good enough. All-bank refresh blocks every bank in the entire chip for the full $t_{RFC}$, which at today's densities is long enough to visibly stall whatever workload is waiting. The fix that DDR5 and LPDDR5 standardized is to split the refresh command itself across banks, so only one bank is ever unavailable at a time while its neighbors keep serving real traffic.

Per-Bank Refresh — Staggering the Outage Instead of Sharing It only one bank is ever unavailable — its neighbors keep serving real traffic the whole time time → bank 0 bank 1 bank 2 bank 3 refreshing (unavailable) available for reads and writes all-bank refresh would gray out every row at once — here at most one bank is ever dark the array still pays the exact same total t_RFC per row over time — staggering it across banks doesn't shrink the tax, it just stops the tax from stalling the whole chip at once

That trade — the same total refresh cost, just spread so it never fully blocks the chip — is the honest summary of everything a refresh controller can do. It has no lever over $t_{RFC}$ itself; that number is set by how much array has to be internally cycled per command, which traces straight back to row count and the capacitor, trench, and word-line physics this project has already covered. Scheduling, bank-splitting, and temperature compensation are all just different ways of paying an unavoidable bill without stalling the rest of the system any more than necessary.

Read the refresh controller through a *scheduling-not-fixing* lens rather than a *"just a background timer"* lens: it inherits $t_{REFW}$ from the 1T1C cell's own $V(t)=V_0e^{-t/\tau}$, it inherits $t_{RFC}$ from how much physical array a refresh command has to touch, and the only real design freedom left — $t_{REFI}=t_{REFW}/N$, the temperature-doubling multiplier, and whether the outage is spread across banks or dumped on the whole chip at once — is about *when* the bill gets paid, never about making the bill smaller. Every keyword this series has covered, from the storage capacitor to the sense amplifier, was about delaying or detecting a leak; the refresh controller is the one component whose entire purpose is simply making sure the bill always gets paid in time.

Take DRAM refresh controller further

Ask the copilot about this term, or have our engineers assess it against your process.