HBM-Pim
Samsung's HBM-PIM places a programmable compute unit inside each bank group of an HBM2 Aquabolt stack, turning the memory itself into a SIMD accelerator that executes FP16 multiply-accumulate directly where the data lives — the first processing-in-memory design to ship as a real, benchmarked product rather than a research prototype. Where the general PIM taxonomy splits into analog crossbars and digital near-memory logic, HBM-PIM is a concrete point in that second category: a production HBM2 die with compute grafted onto its bank structure, not a new memory technology.
What the PCU actually is. Each bank group — a pair of banks that already share peripheral circuitry in a standard DRAM floorplan — gets a programmable computing unit: a 16-wide FP16 multiply-accumulate SIMD engine clocked around 300 MHz, built from logic added to an otherwise conventional HBM2 Aquabolt die. A mode register switches a bank group between ordinary memory behavior and PIM behavior; in PIM mode the PCU reads directly from the bank's row buffer and writes results back, instead of the data making the usual trip up through the TSVs and out across the external interface. The host still talks to the stack over the standard HBM2 JEDEC interface — it issues commands rather than streaming full operands, which is why the change could reach silicon with a driver update rather than a new instruction set.
Why it bandwidth-hacks rather than bandwidth-adds. External HBM bandwidth is already the scarcest resource in an AI system, which is exactly what makes PIM attractive: internal bank bandwidth is far larger than what any external interface can expose, because it never has to cross a TSV, an interposer, or a package boundary. HBM-PIM spends that otherwise-stranded internal bandwidth on FP16 MACs instead of leaving it idle. Samsung's published results — roughly 1.2 TFLOPS per stack, more than double system-level performance, and over 70% lower energy on inference, speech, and recommendation workloads, demonstrated with a Xilinx FPGA standing in for the GPU host — are a direct consequence of avoiding that external round trip rather than of any new transistor-level trick.
Where it sits relative to HMC and the general PIM taxonomy. HMC's logic base die made the memory stack itself a smart, routable device with its own serial interface, discontinued before volume adoption. The general PIM taxonomy splits compute-near-memory into analog crossbars, which perform parallel matrix-vector multiplication as a physical side effect of Ohm's and Kirchhoff's laws, and digital PIM, which adds conventional logic near conventional memory arrays. HBM-PIM is the digital branch's most concrete realization to date: no analog crossbar, no new memory cell, just MAC logic wired into an existing bank structure and gated by a mode bit. It stayed a niche win rather than becoming the default HBM architecture because its MAC-shaped operator set and mode-switch overhead only pay off for a specific slice of workloads — but it is the proof that processing-in-memory can leave a research paper and ship as silicon.