CoWoS & 3D Packaging
CoWoS (Chip-on-Wafer-on-Substrate) and 3D heterogeneous integration are the advanced packaging technologies that make modern AI superchips possible. The semiconductor industry has reached a hard physical ceiling: an extreme ultraviolet (EUV) photolithography scanner can expose an area no larger than a single reticle field of approximately $858\text{ mm}^2$ (roughly $26 \times 33\text{ mm}$). Modern frontier AI accelerators like NVIDIA Blackwell B200 and AMD Instinct MI300X require thousands of square millimeters of compute silicon paired with eight or more High Bandwidth Memory (HBM) stacks. Because monolithic dies of that scale cannot be manufactured with economic yield, foundries assemble multiple modular chiplets side-by-side on an ultra-dense interposer using TSMC's CoWoS platform, and vertically stack active logic dies using 3D bumpless hybrid bonding (TSMC SoIC).
Why interposer density dictates AI throughput. Monolithic printed circuit boards (PCBs) and standard organic package substrates have wire pitch limits of 25 to 50 µm, creating high parasitic capacitance, massive transmission loss, and limited pin density. Placing HBM next to a GPU requires tens of thousands of parallel connections running at multi-gigabit speeds. The interposer acts as a ultra-fine silicon or redistribution bridge with sub-micron line/space (L/S) traces, providing massive interconnect density:
Standard flip-chip packaging with a 130 µm bump pitch yields roughly $59\text{ bumps/mm}^2$. CoWoS microbumps with a 40 µm pitch reach $625\text{ bumps/mm}^2$ — a $10\times$ density multiplier. Moving to bumpless 3D hybrid bonding (Cu-Cu direct contact) with a 0.5 to 1.0 µm pitch shatters this barrier, reaching over $1,000,000\text{ connections/mm}^2$, eliminating solder bump inductance, and dropping communication energy below 0.1 picojoules per bit.
The three variants of TSMC CoWoS:
1. CoWoS-S (Silicon Interposer): The foundational 2.5D architecture powering NVIDIA A100 and H100. Active logic dies and HBM stacks are mounted via microbumps onto a massive, passive silicon wafer containing Through-Silicon Vias (TSVs) and fine metal layers (<0.4 µm L/S). While offering peerless electrical performance and thermal conductivity, fabricating silicon interposers beyond 3.3× reticle size requires complex reticle stitching and carries high wafer cost.
2. CoWoS-R (Organic RDL Interposer): Replaces the rigid silicon interposer with flexible organic thin-film Redistribution Layers (RDL). CoWoS-R reduces cost and improves reliability under thermal cycling for mid-range applications, but is limited to roughly 4 µm L/S routing pitch and suffers higher thermal resistance.
3. CoWoS-L (Local Silicon Interconnect Bridge): The flagship architecture chosen for NVIDIA Blackwell B200. CoWoS-L embeds miniature, high-density local silicon interconnect (LSI) bridges inside a cost-effective, large-area organic RDL substrate. The silicon bridges provide sub-micron dense routing exclusively where dies communicate (GPU-to-GPU and GPU-to-HBM), while the surrounding RDL handles power and ground. This decouples interposer size from the silicon reticle limit, enabling gargantuan interposers exceeding 3.3× to 5.5× reticle fields (over $3,000\text{ mm}^2$).
The vertical leap: 3D SoIC and bumpless hybrid bonding. While CoWoS arranges chiplets horizontally (2.5D), TSMC SoIC (System-on-Integrated-Chips) stacks silicon vertically (3D). Instead of microbumps and solder flux, wafers or dies are polished to atomic flatness via CMP and bonded at room temperature via dielectric ($SiO_2$) hydrophilic contact followed by a thermal anneal that fuses direct copper-to-copper pads. 3D SoIC eliminates the 10–20 µm solder gap, shortens interconnect length to less than 1 µm, cuts latency to the sub-nanosecond domain, and allows AMD MI300X to stack 8 compute dies directly on top of 4 base I/O dies.
| Metric | Flip-Chip BGA | CoWoS-R (RDL) | CoWoS-S (Silicon) | CoWoS-L (Bridge) | 3D SoIC (Hybrid) |
|---|---|---|---|---|---|
| Interconnect Pitch | 100–150 µm | 40–55 µm | 25–40 µm | 25–40 µm (bridge) | <1 to 4 µm |
| Routing Line / Space | >10 µm | 2–4 µm | 0.4–0.8 µm | 0.4–0.8 µm | Backside sub-micron |
| Max Package Size | Single die (~800 mm²) | Up to 2× reticle | Up to 3.3× reticle | >3.3× to 5.5× reticle | 3D Vertical Stack |
| Interconnect Medium | Organic Substrate | Polymer RDL | Passive TSV Silicon | Silicon LSI + RDL | Direct Cu-Cu Bond |
| Energy per Bit | 2–5 pJ/bit | ~1.5 pJ/bit | 0.5–1.0 pJ/bit | 0.4–0.8 pJ/bit | <0.1 pJ/bit |
| HBM Stack Support | None (discrete) | 2–4 stacks | 4–6 stacks | 8+ stacks (HBM3E) | Direct vertical or bus |
| Thermal Warpage Risk | Low | Moderate | Low (matched CTE) | Engineered RDL | Very low (die-to-die) |
| Typical Deployment | Consumer CPUs | Edge AI / Networking | NVIDIA A100 / H100 | NVIDIA Blackwell B200 | AMD MI300 / Apple M-series |