CoWoS & 3D Packaging

CoWoS (Chip-on-Wafer-on-Substrate) and 3D heterogeneous integration are the advanced packaging technologies that make modern AI superchips possible. The semiconductor industry has reached a hard physical ceiling: an extreme ultraviolet (EUV) photolithography scanner can expose an area no larger than a single reticle field of approximately $858\text{ mm}^2$ (roughly $26 \times 33\text{ mm}$). Modern frontier AI accelerators like NVIDIA Blackwell B200 and AMD Instinct MI300X require thousands of square millimeters of compute silicon paired with eight or more High Bandwidth Memory (HBM) stacks. Because monolithic dies of that scale cannot be manufactured with economic yield, foundries assemble multiple modular chiplets side-by-side on an ultra-dense interposer using TSMC's CoWoS platform, and vertically stack active logic dies using 3D bumpless hybrid bonding (TSMC SoIC).

Why interposer density dictates AI throughput. Monolithic printed circuit boards (PCBs) and standard organic package substrates have wire pitch limits of 25 to 50 µm, creating high parasitic capacitance, massive transmission loss, and limited pin density. Placing HBM next to a GPU requires tens of thousands of parallel connections running at multi-gigabit speeds. The interposer acts as a ultra-fine silicon or redistribution bridge with sub-micron line/space (L/S) traces, providing massive interconnect density:

$$ D_{interconnect} = \frac{1}{P_{bump}^2} \quad \left[\text{interconnects / mm}^2\right] $$

Standard flip-chip packaging with a 130 µm bump pitch yields roughly $59\text{ bumps/mm}^2$. CoWoS microbumps with a 40 µm pitch reach $625\text{ bumps/mm}^2$ — a $10\times$ density multiplier. Moving to bumpless 3D hybrid bonding (Cu-Cu direct contact) with a 0.5 to 1.0 µm pitch shatters this barrier, reaching over $1,000,000\text{ connections/mm}^2$, eliminating solder bump inductance, and dropping communication energy below 0.1 picojoules per bit.

The three variants of TSMC CoWoS:
1. CoWoS-S (Silicon Interposer): The foundational 2.5D architecture powering NVIDIA A100 and H100. Active logic dies and HBM stacks are mounted via microbumps onto a massive, passive silicon wafer containing Through-Silicon Vias (TSVs) and fine metal layers (<0.4 &#181;m L/S). While offering peerless electrical performance and thermal conductivity, fabricating silicon interposers beyond 3.3&#215; reticle size requires complex reticle stitching and carries high wafer cost.
2. CoWoS-R (Organic RDL Interposer): Replaces the rigid silicon interposer with flexible organic thin-film Redistribution Layers (RDL). CoWoS-R reduces cost and improves reliability under thermal cycling for mid-range applications, but is limited to roughly 4 &#181;m L/S routing pitch and suffers higher thermal resistance.
3. CoWoS-L (Local Silicon Interconnect Bridge): The flagship architecture chosen for NVIDIA Blackwell B200. CoWoS-L embeds miniature, high-density local silicon interconnect (LSI) bridges inside a cost-effective, large-area organic RDL substrate. The silicon bridges provide sub-micron dense routing exclusively where dies communicate (GPU-to-GPU and GPU-to-HBM), while the surrounding RDL handles power and ground. This decouples interposer size from the silicon reticle limit, enabling gargantuan interposers exceeding 3.3&#215; to 5.5&#215; reticle fields (over $3,000\text{ mm}^2$).

The vertical leap: 3D SoIC and bumpless hybrid bonding. While CoWoS arranges chiplets horizontally (2.5D), TSMC SoIC (System-on-Integrated-Chips) stacks silicon vertically (3D). Instead of microbumps and solder flux, wafers or dies are polished to atomic flatness via CMP and bonded at room temperature via dielectric ($SiO_2$) hydrophilic contact followed by a thermal anneal that fuses direct copper-to-copper pads. 3D SoIC eliminates the 10–20 &#181;m solder gap, shortens interconnect length to less than 1 &#181;m, cuts latency to the sub-nanosecond domain, and allows AMD MI300X to stack 8 compute dies directly on top of 4 base I/O dies.

MetricFlip-Chip BGACoWoS-R (RDL)CoWoS-S (Silicon)CoWoS-L (Bridge)3D SoIC (Hybrid)
Interconnect Pitch100&#8211;150 &#181;m40&#8211;55 &#181;m25&#8211;40 &#181;m25&#8211;40 &#181;m (bridge)&lt;1 to 4 &#181;m
Routing Line / Space&gt;10 &#181;m2&#8211;4 &#181;m0.4&#8211;0.8 &#181;m0.4&#8211;0.8 &#181;mBackside sub-micron
Max Package SizeSingle die (~800 mm&#178;)Up to 2&#215; reticleUp to 3.3&#215; reticle&gt;3.3&#215; to 5.5&#215; reticle3D Vertical Stack
Interconnect MediumOrganic SubstratePolymer RDLPassive TSV SiliconSilicon LSI + RDLDirect Cu-Cu Bond
Energy per Bit2&#8211;5 pJ/bit~1.5 pJ/bit0.5&#8211;1.0 pJ/bit0.4&#8211;0.8 pJ/bit&lt;0.1 pJ/bit
HBM Stack SupportNone (discrete)2&#8211;4 stacks4&#8211;6 stacks8+ stacks (HBM3E)Direct vertical or bus
Thermal Warpage RiskLowModerateLow (matched CTE)Engineered RDLVery low (die-to-die)
Typical DeploymentConsumer CPUsEdge AI / NetworkingNVIDIA A100 / H100NVIDIA Blackwell B200AMD MI300 / Apple M-series
CoWoS & 3D Packaging Architecture Chip-on-Wafer-on-Substrate interconnects GPU compute dies with HBM3E memory stacks across a multi-reticle interposer HBM3E Stack 1 12-Hi / 16-Hi DRAM TSVs in DRAM AI Compute Die 1 Sub-2nm GPU / Accelerator Tensor Core Logic + Cache AI Compute Die 2 Sub-2nm GPU / Accelerator Tensor Core Logic + Cache HBM3E Stack 2 12-Hi / 16-Hi DRAM TSVs in DRAM FINE-PITCH MICROBUMPS (25–40 µm PITCH) OR 3D SoIC Cu-Cu HYBRID BONDS (<1 µm) CoWoS-L / CoWoS-S INTERPOSER (>3.3× RETICLE FIELD) LSI Bridge Die-Die LSI Bridge Sub-micron L/S Interposer Traces + TSVs Sub-micron L/S Interposer Traces + TSVs C4 Bumps (100–150 µm pitch) Organic Package Substrate & BGA Balls to Server Motherboard The Engine of Modern AI Accelerators Reticle Limit Ceiling Monolithic: ~858 mm² CoWoS-L: >3,300 mm² Interconnect Density CoWoS Microbumps: 625/mm² 3D SoIC: >1,000,000/mm² Energy Efficiency PCB Traces: ~5 pJ/bit CoWoS/SoIC: <0.5 pJ/bit HBM Memory Width 8 Stacks · 8,192-bit Up to 24 TB/s throughput

Take cowos & 3d packaging further

Ask the copilot about this term, or have our engineers assess it against your process.