Home Knowledge Base Parallel Matrix Multiplication (GEMM)

Parallel Matrix Multiplication (GEMM) is the operation C = α×A×B + β×C for dense matrices — the most important computational kernel in high-performance computing, deep learning, and scientific simulations, extensively optimized in BLAS libraries.

GEMM Complexity

BLAS (Basic Linear Algebra Subprograms)

CPU GEMM Optimization Techniques

Blocking (Tiling):

SIMD Vectorization:

Packing:

GPU GEMM

Distribution Across Nodes

GEMM optimization is the foundation of modern AI computing — 70–90% of transformer inference and training time is spent in matrix multiplications, and every 10% GEMM efficiency improvement translates directly to training cost and inference latency reduction.

parallel matrix multiplicationgemmblas libraryhigh performance matrixmatmul optimization

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.