Upmem

UPMEM builds a PIM-enabled DIMM that is pin-compatible with a standard DDR4 slot, so it plugs into an ordinary server motherboard rather than requiring an interposer, a new memory protocol, or a graphics-card socket — the most compatible point yet in the family of PIM designs that includes HBM-PIM and GDDR6-AiM. The difference inside the chip is just as deliberate: instead of a fixed-function MAC unit, each UPMEM DRAM Processing Unit (DPU) is a full general-purpose 32-bit RISC core, trading peak arithmetic density for the ability to run arbitrary, branching, pointer-chasing code next to its own private bank of DRAM.

UPMEM: a General-Purpose RISC Core Per DRAM Bank a standard DDR4 DIMM slot — programmability traded for peak arithmetic density PIM-Enabled DIMM — Standard DDR4 Form Factor Same DDR4 edge connector, same slot PIM chips across the module — each chip holds several DPUs One DPU — General-Purpose, Not Fixed-Function 32-bit RISC core IRAM ~24 KB WRAM ~64 KB Deeply multithreaded — many hardware threads round-robin to hide DRAM access latency instead of relying on a large cache. exclusive access to its own ~64 MB DRAM bank What a RISC Core Buys and Costs Where DPUs win Scan/filter, graph traversal, genome string matching — branchy, low-arithmetic-intensity, pointer-chasing work that a fixed MAC array cannot execute at all. Thousands of DPUs run independent program counters, each against its own bank, in parallel. Where DPUs lose Dense matrix multiply and large neural networks need shared caches, tensor cores, and far higher peak FLOPs than a scalar RISC core per bank can deliver. a MAC array spends on nothing but multipliers. The compatibility spectrum, completed HMC: new stack, new interface — discontinued. HBM-PIM: HBM2 stack, new protocol, FPGA host demo. AiM: GDDR6 pins, graphics-card socket. UPMEM: DDR4 pins, standard server DIMM slot.

Programming a few thousand independent cores. A PIM-enabled server exposes on the order of thousands of DPUs across its DIMMs — a figure UPMEM quotes at roughly 2,560 DPUs in a fully populated system — each running its own program and owning its own DRAM bank exclusively, so there is no bank-contention problem to solve: parallelism comes for free from the hardware layout, and the work is dispatching independent tasks to independent cores rather than extracting SIMD lanes from a single shared array. UPMEM ships an SDK so a host application can compile C for the DPU instruction set, hand each DPU its slice of data, and collect results — a conventional multithreaded-programming model applied to memory banks, not a new numerical abstraction like a MAC command or a crossbar conductance map.

Why this is the opposite trade from AiM and HBM-PIM. AiM and HBM-PIM both chose fixed-function MAC units: narrow, dense, and fast per unit of area, but limited to the operations their designers wired in. UPMEM chose a real instruction set: flexible enough to run scan, filter, sort, string-matching, or graph code that a MAC array simply cannot express, at the cost of lower peak arithmetic throughput per die. Published speedups cluster in the single-digit-to-tens-of-times range for scan-and-filter, graph, and genomic-search workloads — large relative to moving the same data across a conventional memory bus, modest next to a GPU's raw FLOPs on workloads the GPU is actually suited for. UPMEM is the end of this family's spectrum where programmability, not peak density, is the entire point: the same DDR4 slot any server already has, filled with a part that runs software instead of one fixed operation.

Take UPMEM further

Ask the copilot about this term, or have our engineers assess it against your process.