DRAM Buried Word Line

# DRAM Buried Word Line: Saddle-Fin Gate Architecture, the GIDL Trade-Off, and Hierarchical Row Activation

A buried word line (BWL) is a DRAM access-transistor gate that no longer sits on top of the silicon — it is recessed into a trench etched below the original surface, wrapping around a narrow silicon fin on three sides instead of controlling the channel from one flat face above it. DRAM moved to this geometry because a 1T1C access transistor at DRAM's pitch cannot get a planar gate long enough to stay off reliably; recessing the gate buys channel length without spending any extra footprint, at the cost of a leakage trade-off that did not exist when the gate sat flat on top of the silicon.

Saddle-Fin RCAT — Why the Word Line Dips Below the Surface wrapping the gate around a silicon fin adds channel length without adding footprint p-type silicon substrate original surface source drain GIDL hotspot — field crowds where gate meets drain V_WL = V_pp (boosted above V_dd + V_th so the full V_dd passes) the gate wraps the fin on three sides — left wall, bottom, right wall — a dimension this flat cross-section can't fully show, which is why it's called a "saddle": it sits astride the fin the way a saddle sits a horse off-state, the same word line gets driven negative — which helps one leakage path and hurts another

Getting a full logic voltage onto the storage node at all requires overdriving the gate. An NMOS access transistor can only pass a voltage up to $V_{dd} - V_{th}$ onto its source before it starts to turn itself off, so writing a full $V_{dd}$ onto the storage capacitor needs the word line driven above the supply rail entirely:

$$ V_{WL} \geq V_{dd} + V_{th} + V_{margin} $$

Every DRAM chip generates this boosted voltage, called $V_{pp}$, on-die with a dedicated charge pump, purely so the access transistor can pass a clean "1" without losing a threshold voltage's worth of charge on the way in.

That same transistor has to shut off just as cleanly, and that is where the recessed geometry creates a new problem. A buried gate's sidewall sits close to, and slightly underlaps, the drain diffusion above it — exactly the geometry that produces gate-induced drain leakage (GIDL): band-to-band tunneling at the gate-drain overlap, driven by the electric field there, which gets *stronger*, not weaker, the more negative the word line is pushed in the off state. Subthreshold leakage through the channel and GIDL at the drain overlap respond to the off-state bias in opposite directions, which makes the word line's "off" voltage a genuine optimization, not just "as negative as possible":

The Off-State Trade-Off — Subthreshold Leakage vs. GIDL driving the word line more negative fixes one leakage path and opens another more negative off-state bias → leakage current (log scale) → subthreshold leakage — falls GIDL — rises optimal V_kk — the array's actual back-bias setpoint I_total(V) = I_subthreshold(V) + I_GIDL(V) — minimized away from both extremes, not at either one this single minimum is why the off-state word-line voltage is a tuned chip parameter, not just "as negative as possible"

That minimum is tuned once per design, but the word line still has to physically reach every cell fast enough, and a buried line is a worse conductor than a surface metal one. Recessed polysilicon or tungsten inside a narrow trench carries far more resistance per micron than a metal interconnect line, and the RC delay of driving it from one end scales with the square of its length:

$$ t_{WL} \propto R_{WL} \cdot C_{WL} \cdot N^2 $$

Doubling how many cells one continuous word line serves quadruples its activation delay — an unacceptable trade at DRAM's density — so no production array drives a buried word line end-to-end from a single driver.

Why One Row Activation Needs Hundreds of Local Drivers a buried word line is too resistive to drive the whole array from one end master word line (metal, low resistance) — spans the whole array SWL driver SWL driver SWL driver local buried word-line segments — high resistance, kept short on purpose each driver activates only its own few hundred cells — never the full row in one hop t_WL ∝ R_WL · C_WL · N² — doubling one segment's cell count quadruples its delay, which is why hierarchical master/sub-word-line splitting, not a faster driver, is the actual fix

All three constraints — the $V_{pp}$ overdrive, the GIDL-versus-subthreshold minimum, and the hierarchical driver split — are decided together, not separately. Recess depth and gate work function set how much overdrive the transistor needs and how sharp the GIDL corner is; the chosen $V_{kk}$ back-bias sets how hard the charge pumps and substrate-bias generators on the die have to work; and the sub-word-line pitch sets how many drivers a given array floorplan needs room for. A DRAM maker that designs the cell, the process recipe, and the on-die voltage generation together can tune all three against each other; a team trying to license just the transistor geometry inherits none of that co-design, which is the same reason the capacitor and the isolation trench turned out to be moats in their own right — the buried word line is a third structure in the same 1T1C cell where the manufacturing process and the circuit design are inseparable.

Read the buried word line through a *field-at-the-corner* lens rather than a *"just a recessed gate"* lens: every design choice here — how deep to recess it, how negative to drive it off, how many cells one segment gets to serve — is a different point on the same curve in $I_{total}(V) = I_{subthreshold}(V) + I_{GIDL}(V)$, traded off against the $t_{WL} \propto R_{WL}C_{WL}N^2$ delay that decides how finely the array has to be split to activate a row fast enough to matter.

Take DRAM buried word line further

Ask the copilot about this term, or have our engineers assess it against your process.