HBM-Pim

Samsung's HBM-PIM places a programmable compute unit inside each bank group of an HBM2 Aquabolt stack, turning the memory itself into a SIMD accelerator that executes FP16 multiply-accumulate directly where the data lives — the first processing-in-memory design to ship as a real, benchmarked product rather than a research prototype. Where the general PIM taxonomy splits into analog crossbars and digital near-memory logic, HBM-PIM is a concrete point in that second category: a production HBM2 die with compute grafted onto its bank structure, not a new memory technology.

HBM-PIM: a Programmable Compute Unit Inside Each Bank Group built on HBM2 Aquabolt — the PCU reuses internal bank bandwidth the external pins never see One DRAM Die, Two Bank Groups Bank Group A bank 0 — normal mode bank 1 — normal mode PCU FP16 MAC · 300 MHz shared across banks Bank Group B bank 0 — normal mode bank 1 — normal mode PCU FP16 MAC · 300 MHz shared across banks A bank group is either a memory bank or a compute unit — a mode register toggles it, no new pins, no new protocol layer. Why this uses bandwidth the host never sees A row-buffer read inside the bank feeds the PCU directly — no trip across the TSVs, interposer, or external HBM channel. Internal bank bandwidth vastly exceeds the ~1.2 TB/s the pins expose, so PIM mode spends bandwidth the interface can't. Host Command Flow (Unmodified HBM Interface) GPU / FPGA host standard JEDEC command base logic die PCU reads row buffer, MACs reduced result returns Benchmarked (Samsung, 2021) 1.2 TFLOPS/stack (FP16) >2× system speedup >70% lower energy on AI inference, speech, recommendation kernels Demonstrated on a Xilinx Virtex UltraScale+ FPGA in place of a GPU host Driver-level software change — no new host ISA required Why it stayed a niche win, not the default Narrow operator set (MAC-shaped kernels only), mode-switch overhead, and limited precision.

What the PCU actually is. Each bank group — a pair of banks that already share peripheral circuitry in a standard DRAM floorplan — gets a programmable computing unit: a 16-wide FP16 multiply-accumulate SIMD engine clocked around 300 MHz, built from logic added to an otherwise conventional HBM2 Aquabolt die. A mode register switches a bank group between ordinary memory behavior and PIM behavior; in PIM mode the PCU reads directly from the bank's row buffer and writes results back, instead of the data making the usual trip up through the TSVs and out across the external interface. The host still talks to the stack over the standard HBM2 JEDEC interface — it issues commands rather than streaming full operands, which is why the change could reach silicon with a driver update rather than a new instruction set.

Why it bandwidth-hacks rather than bandwidth-adds. External HBM bandwidth is already the scarcest resource in an AI system, which is exactly what makes PIM attractive: internal bank bandwidth is far larger than what any external interface can expose, because it never has to cross a TSV, an interposer, or a package boundary. HBM-PIM spends that otherwise-stranded internal bandwidth on FP16 MACs instead of leaving it idle. Samsung's published results — roughly 1.2 TFLOPS per stack, more than double system-level performance, and over 70% lower energy on inference, speech, and recommendation workloads, demonstrated with a Xilinx FPGA standing in for the GPU host — are a direct consequence of avoiding that external round trip rather than of any new transistor-level trick.

Where it sits relative to HMC and the general PIM taxonomy. HMC's logic base die made the memory stack itself a smart, routable device with its own serial interface, discontinued before volume adoption. The general PIM taxonomy splits compute-near-memory into analog crossbars, which perform parallel matrix-vector multiplication as a physical side effect of Ohm's and Kirchhoff's laws, and digital PIM, which adds conventional logic near conventional memory arrays. HBM-PIM is the digital branch's most concrete realization to date: no analog crossbar, no new memory cell, just MAC logic wired into an existing bank structure and gated by a mode bit. It stayed a niche win rather than becoming the default HBM architecture because its MAC-shaped operator set and mode-switch overhead only pay off for a specific slice of workloads — but it is the proof that processing-in-memory can leave a research paper and ship as silicon.

Take HBM-PIM further

Ask the copilot about this term, or have our engineers assess it against your process.