Home Knowledge Base GPU (Graphics Processing Unit)

GPU (Graphics Processing Unit) is the massively parallel processor that has become the primary hardware accelerator for deep learning — containing thousands of cores optimized for the matrix multiplications and tensor operations that dominate neural network training and inference, delivering 10-100x speedups over CPUs and fundamentally enabling the modern AI revolution from transformer models to generative AI through sheer computational throughput and high-bandwidth memory architectures.

What Is a GPU?

Modern AI GPU Architecture

ComponentPurposeExample (H100)
CUDA CoresGeneral-purpose parallel computation16,896 cores
Tensor CoresSpecialized matrix multiply-accumulate units528 (4th gen)
HBM (High Bandwidth Memory)High-speed memory for model weights and activations80GB HBM3 at 3.35 TB/s
NVLinkHigh-bandwidth GPU-to-GPU interconnect900 GB/s bidirectional
Transformer EngineAutomatic mixed-precision for transformersFP8 support

Key NVIDIA GPU Generations for AI

Why GPUs Matter for AI

GPU Programming Ecosystem

Cloud GPU Access

GPUs are the engine powering the entire modern AI revolution — providing the massive parallel compute throughput that makes training billion-parameter models feasible and inference at scale affordable, with GPU supply and innovation directly determining the pace of AI progress worldwide.

gpu (graphics processing unit)graphics processing unithardware

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.