Smart Ssd
Computational storage embeds compute — typically an FPGA or a small ARM core cluster — inside the SSD itself, between the NAND flash array and the host's NVMe/PCIe interface, so operations like filtering, compression, or decompression happen in-line on data that would otherwise have to cross the bus untouched just to be reduced on the host. Samsung's SmartSSD is the best-known commercial example: a standard U.2 NVMe drive with an embedded FPGA that runs user-programmable kernels on data as it streams off the flash. Where HBM-PIM, AiM, and UPMEM all push compute closer to DRAM, computational storage pushes it one layer further out — all the way to the nonvolatile tier, where the bottleneck is not DRAM or HBM bandwidth but the comparatively narrow, comparatively slow NVMe link itself.
What actually moves, and what stays the same. The NAND array, the flash translation layer, and the NAND controller's wear-leveling and ECC duties are untouched — a computational storage drive is still, first, a normal SSD. What changes is that an FPGA or small processor cluster sits in the data path between the flash controller and the NVMe interface, able to run a user-programmable kernel on data before it leaves the drive. A database scan that would normally read every row across PCIe only to discard 90% of them on the host can instead apply the filter inside the drive and send back only the matching rows; a drive doing transparent compression, as ScaleFlux's CSD line does, writes less to NAND and returns decompressed data to the host without the host CPU ever running the codec.
Why the SSD, and not just a bigger host CPU budget. NVMe bandwidth, while far ahead of older storage interfaces, remains well below DRAM or HBM bandwidth, and every byte that crosses it to be immediately discarded is a byte that bandwidth, host CPU cycles, and power were spent moving for nothing. Computational storage is the same argument that motivates HBM-PIM, AiM, and UPMEM — move the computation to the data instead of the data to the computation — applied at the layer where the data is biggest, cheapest, and most durable, but also furthest from the host and slowest to reach. It is the natural end of the "close to the data" spectrum this family of designs traces: DRAM bank, HBM stack, GDDR6 die, DDR4 DIMM, and now the NAND array itself, each tier trading proximity and bandwidth for capacity and persistence.