Hbm4 Yield
# HBM4 Yield and Known-Good-Die Testing: Compound Yield Economics, Pre-Bond Test Architecture, and Why KGD Is Non-Negotiable
Once a die is hybrid-bonded into an HBM4 stack, it can no longer be replaced. There is no rework step that un-bonds a single bad DRAM die out of a twelve-to-sixteen-high stack without destroying every good die bonded alongside it. That one fact is why known-good-die (KGD) testing — proving every die is good *before* it is ever bonded — is not a quality nicety for HBM4, it is the only point in the entire process where a defect can still be discarded cheaply.
The reason this gate has to be airtight is simple multiplication. If every die in a stack has an independent probability $Y_{die}$ of being good, the probability that an entire $N$-die stack is good is the product of all of them:
That exponent is brutal. A die yield that sounds comfortable on its own — say 99.5% — still only clears a 16-high stack 92% of the time; at 98% per die, a 16-high stack is good barely 72% of the time. HBM4's stacks are taller than any prior generation, so the same per-die yield that was tolerable for an 8-high HBM3E stack can become a serious problem once raised to the 16th power.
That same multiplication is also why catching a bad die gets dramatically more expensive the later it is caught. A die that fails at wafer probe costs exactly one die. The same defect, if it slips through pre-bond screening and is only caught after bonding, costs every good die stacked with it, the bonding process itself, and — for HBM4 specifically — the logic-node base die the base-die article already established is the single most expensive component in the assembly. Caught only after the module ships, it costs a field failure on top of all of that.
That cost curve is the actual economic justification for a test step that is, by itself, harder and more expensive than testing an ordinary packaged part. Probing a bare die at the wafer level means contacting pads that were never designed to survive repeated mechanical probing, running the base die's full PHY and BIST at speed without any of the packaging that normally helps with power delivery and signal integrity, and in many flows adding a wafer-level burn-in step specifically to catch the latent defects that would otherwise only show up as field failures. None of that would be worth doing if a bad die discovered later were only as expensive as the die itself — it is worth doing precisely because $Y_{stack}=Y_{die}^N$ turns one missed defect into the cost of everything bonded around it.
Read HBM4 yield through a *irreversibility* lens rather than a *"more testing is always better"* lens: the hybrid-bonding step this project already covered is exactly the point after which a bad die can no longer be removed, so every dollar spent on pre-bond known-good-die testing is bought back many times over by what $Y_{stack}=Y_{die}^N$ would otherwise cost once that same defect is sealed inside a finished stack.