Aim Accelerator in Memory

SK hynix's AiM (Accelerator-in-Memory) puts MAC units inside a GDDR6 die's global buffer rather than inside an HBM stack, and keeps the standard GDDR6 pinout unchanged — so a GDDR6-AiM chip can sit in the same socket a conventional graphics-memory controller expects, instead of requiring an interposer, a new protocol, or a redesigned host. Where Samsung's HBM-PIM proved bank-level compute could ship inside an HBM2 stack, AiM proves the same idea works in commodity GDDR6, trading HBM's higher bandwidth ceiling for a part that drops into existing graphics-memory infrastructure.

GDDR6-AiM: MAC Units in the Global Buffer, Same Pinout a few extra commands on the existing GDDR6 bus — no interposer, no new socket GDDR6-AiM Die Floorplan 16 Banks (conventional DRAM array) bank group reads feed the global buffer below Global Buffer — 16 MAC Units + Accumulators MAC + accumulate on data already in the buffer — no round trip off-die Standard GDDR6 pinout — unchanged A Few Extra Commands, Same Bus GDDR6 baseline commands still work ACT · RD · WR · REF — unchanged, full-speed memory access Added: MAC, WRGB, RDMAC-style ops Issued on the same command pins — host needs driver support, not a new physical interface. Target: memory-bound AI inference Embedding lookups and GEMV-shaped kernels where bandwidth, not FLOPs, sets the achievable rate — recommendation models and transformer decoding are the published target workloads. Published figures are order-of-magnitude, not a datasheet guarantee. AiMX: Many Dies, One Accelerator Card GDDR6-AiM dies on a PCIe card — a drop-in inference accelerator, not a GPU replacement, built from parts that fit existing sockets.

Why pin compatibility is the point. A GDDR6 controller already exists in every graphics card and many inference accelerators; AiM's bet is that reusing that controller's physical interface — the same pins, the same baseline `ACT`/`RD`/`WR`/`REF` command set — lets the compute ship without asking the rest of the system to change first. The MAC and accumulate logic lives in the die's global buffer, the circuitry that already aggregates bank reads before they leave the chip, so the added commands (`MAC`, buffer-write, result-read variants) ride on the existing command bus rather than demanding new physical pins. The cost of that compatibility is scope: AiM's compute is sized to what a GDDR6 global buffer can host, well short of an HBM stack's internal bandwidth or a GPU's arithmetic throughput.

Where it fits next to HBM-PIM and HMC. HMC built an entirely new stack with its own logic base die and serial interface, and was discontinued. HBM-PIM kept the HBM2 stack but changed its protocol surface enough to need a driver and a non-standard host (a Xilinx FPGA in Samsung's own demo). AiM goes one step further toward compatibility: same memory technology category as the graphics cards already shipping, same pins, an accelerator (AiMX) built by stacking several AiM dies on a card rather than requiring a new socket. Each step trades generality for deployability — HMC changed everything and vanished, HBM-PIM changed the protocol and shipped in benchmarks, AiM changes almost nothing visible to the board and targets the narrowest, most bandwidth-starved slice of AI inference: embedding lookups and GEMV-shaped kernels where moving data, not computing on it, was always the bottleneck.

Take AiM accelerator in memory further

Ask the copilot about this term, or have our engineers assess it against your process.