DRAM Sense Amplifier

# DRAM Sense Amplifier: Cross-Coupled Latch Architecture, Regenerative Feedback Theory, and the Folded-Bitline Application

The sense amplifier is where a DRAM read actually becomes a decision. The 1T1C cell's charge-sharing step leaves the bit line only tens of millivolts away from its precharged midpoint — nowhere close to a clean logic level. Every DRAM sense amplifier is a cross-coupled latch: two inverters wired so each one's output drives the other's input, which does nothing while they agree, but the instant one side is even a few millivolts ahead, that pair of inverters actively fights to make the gap bigger, not smaller, until one side is pinned at the supply rail and the other at ground.

The Cross-Coupled Latch — Two Inverters Racing to a Decision whichever side starts even a few millivolts ahead wins, and positive feedback does the rest EQ precharges both rails to V_dd/2 before every sense BL BL̄ ISO ISO to array → ← to array SAP PMOS_L PMOS_R NMOS_L NMOS_R SAN each gate is driven by the opposite node — that cross-wiring is the entire mechanism firing SAP and SAN connects both cross-coupled pairs to power at once — whichever side started even slightly ahead turns its own pull-up or pull-down on harder, which pulls its partner further behind — a loop that only stops once one rail hits V_dd and the other hits 0

That loop is positive feedback, and positive feedback regenerates exponentially, not linearly. Once SAP and SAN fire, the voltage difference between the two rails grows away from whatever it started at, at a rate set by the latch's own transconductance and the capacitance it has to swing:

$$ \Delta V(t) = \Delta V(0) \cdot e^{t/\tau}, \qquad \tau = \frac{C_{node}}{g_m} $$

The exponent is the whole story: a healthy cell leaves a comfortable $\Delta V(0)$ and the latch resolves it almost immediately, but a cell that leaked more than it should have — through retention, through STI corner leakage, through GIDL at a buried word line — hands the latch a smaller $\Delta V(0)$ to start from, and the exact same exponential takes measurably longer to reach a safe logic level.

Regeneration Is a Race Against Time, Not Just Gain a smaller starting imbalance takes measurably longer to resolve — and can run out of time time since SAP/SAN fire → bit-line voltage difference ΔV → sense window closes healthy cell — resolves early typical cell marginal cell — still rising when time runs out undecided read ΔV(t) = ΔV(0) · e^(t/τ), τ = C_node/g_m — every upstream leakage path shrinks ΔV(0) and borrows from this margin

That is the direct link back to everything upstream of the amplifier. $\Delta V(0)$ is exactly the charge-sharing result from the 1T1C cell — $\Delta V(0) = V_{cell} \cdot C_s/(C_s + C_{BL})$ — so a slightly leaky trench, a slightly worse GIDL corner, or simply waiting a few milliseconds too long past the refresh deadline all show up here as the same thing: a smaller number handed to the exponential, and less margin before the sense window closes. The sense amplifier doesn't create read errors on its own; it inherits and then race-condition-amplifies whatever margin survived everything upstream.

That exponential has a second, quieter enemy: the two halves of the latch are never perfectly identical. Random dopant variation gives the two matched transistors in each cross-coupled pair slightly different threshold voltages, which biases the latch toward one side before any cell signal even arrives. That mismatch follows Pelgrom's law — it shrinks only with the square root of transistor area, not with area itself:

$$ \sigma(\Delta V_{th}) = \frac{A_{V_{th}}}{\sqrt{W \cdot L}} $$

Every node that shrinks the sense transistors to save area pays for it with worse matching, at the same time the storage capacitor is handing the latch a smaller signal to resolve — the two curves this industry has spent decades fighting both move the wrong way together.

One Sense Amp, Shared by Two Half-Arrays folding the bit line lets neighboring half-arrays share one amplifier and cancel common noise top half-array active this cycle — carries the signal bottom half-array idle this cycle — holds the quiet reference BL (true) ISO_top shared sense amp BL̄ (complement) ISO_bot only one half-array drives real data per cycle — the other supplies the quiet V_dd/2 reference one amplifier instead of one-per-half-array saves area, and routing the reference through the same physical layout as the signal cancels supply noise that would otherwise look like real ΔV

Folding the bit line this way is why DRAM arrays can afford one sense amplifier for every two half-arrays instead of one per column of active cells. The idle half-array isn't wasted — routing its matched bit line through the same physical layout as the active signal means that any noise riding on the shared supply rails cancels out of the difference the latch actually measures, which is exactly the kind of common-mode rejection a bare single-ended comparison could never get.

Read the DRAM sense amplifier through a *margin-and-mismatch* lens rather than a *"just a comparator"* lens: every millivolt of $\Delta V(0)$ it is handed came from $Q=C_sV_{cell}$ surviving every leakage path upstream, every microsecond it has to resolve that signal is governed by the same $\Delta V(t)=\Delta V(0)e^{t/\tau}$ this project has now traced from the storage capacitor through the isolation trench and the buried word line, and every transistor pair inside it is fighting Pelgrom's $\sigma(\Delta V_{th}) \propto 1/\sqrt{WL}$ just to stay neutral until the real signal arrives. The amplifier is the last link in the chain — it doesn't generate the DRAM scaling problem, it is simply where every upstream compromise finally gets cashed in.

Take DRAM sense amplifier further

Ask the copilot about this term, or have our engineers assess it against your process.