Home Knowledge Base The roofline shows exactly where AMD's memory advantage is not a rounding error.

ROCm is AMD's open-source GPU compute stack: the CUDA-equivalent layer of compiler, runtime and math libraries that turns an Instinct accelerator into a machine-learning device, with HIP for the kernels, rocBLAS and hipBLASLt for matrix multiply, MIOpen and Composable Kernel for the fused operators, and RCCL for the collectives. The hardware underneath it is, on paper, ahead. An AMD MI300X carries 192 GB of HBM3 at 5.3 TB/s against an NVIDIA H100's 80 GB at 3.35 TB/s, and quotes 1,307 dense BF16 TFLOP/s against 989. Almost every number that sells a datacenter GPU favors AMD by 30 to 140 percent. The gap that still keeps CUDA the default is in none of them, and that is the whole story of ROCm.

<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 760 470" font-family="Georgia, serif" font-size="13"><rect x="0" y="0" width="760" height="470" fill="#0d1117"/><text x="380" y="30" text-anchor="middle" font-size="21" font-weight="700" fill="#e6edf3">Where AMD wins on hardware, and where CUDA wins it back</text><text x="380" y="50" text-anchor="middle" font-size="13" fill="#8b98a5">The MI300X&#39;s memory edge is real; the software surface area decides whether it survives</text><rect x="66" y="74" width="290" height="314" fill="#161b22" stroke="#30363d"/><text x="211.0" y="68" text-anchor="middle" font-size="13" font-weight="bold" fill="#c9d1d9">Roofline: attainable BF16 throughput</text><line x1="90.2" y1="74" x2="90.2" y2="388" stroke="#21262d"/><text x="90.2" y="404" text-anchor="middle" font-size="11" fill="#8b98a5">1</text><line x1="170.7" y1="74" x2="170.7" y2="388" stroke="#21262d"/><text x="170.7" y="404" text-anchor="middle" font-size="11" fill="#8b98a5">10</text><line x1="251.3" y1="74" x2="251.3" y2="388" stroke="#21262d"/><text x="251.3" y="404" text-anchor="middle" font-size="11" fill="#8b98a5">100</text><line x1="331.8" y1="74" x2="331.8" y2="388" stroke="#21262d"/><text x="331.8" y="404" text-anchor="middle" font-size="11" fill="#8b98a5">1000</text><line x1="66" y1="388.0" x2="356" y2="388.0" stroke="#21262d"/><text x="60" y="392.0" text-anchor="end" font-size="11" fill="#8b98a5">10</text><line x1="66" y1="251.5" x2="356" y2="251.5" stroke="#21262d"/><text x="60" y="255.5" text-anchor="end" font-size="11" fill="#8b98a5">100</text><line x1="66" y1="115.1" x2="356" y2="115.1" stroke="#21262d"/><text x="60" y="119.1" text-anchor="end" font-size="11" fill="#8b98a5">1000</text><text x="211.0" y="422" text-anchor="middle" font-size="12" fill="#c9d1d9">arithmetic intensity (FLOP / byte)</text><text x="24" y="231.0" text-anchor="middle" font-size="12" fill="#c9d1d9" transform="rotate(-90 24 231.0)">TFLOP/s</text><rect x="90.2" y="74" width="121.2" height="314" fill="#58a6ff" opacity="0.10"/><text x="152.9" y="90" text-anchor="middle" font-size="11" fill="#58a6ff">decode</text><line x1="300.6" y1="74" x2="300.6" y2="388" stroke="#d29922" stroke-dasharray="3,3"/><text x="300.6" y="90" text-anchor="middle" font-size="11" fill="#d29922">prefill</text><path d="M 128.5 388.0 L 133.4 379.7 L 138.2 371.4 L 143.1 363.2 L 148.0 354.9 L 152.9 346.6 L 157.8 338.3 L 162.7 330.0 L 167.6 321.7 L 172.5 313.5 L 177.3 305.2 L 182.2 296.9 L 187.1 288.6 L 192.0 280.3 L 196.9 272.0 L 201.8 263.8 L 206.7 255.5 L 211.5 247.2 L 216.4 238.9 L 221.3 230.6 L 226.2 222.3 L 231.1 214.1 L 236.0 205.8 L 240.9 197.5 L 245.8 189.2 L 250.6 180.9 L 255.5 172.6 L 260.4 164.4 L 265.3 156.1 L 270.2 147.8 L 275.1 139.5 L 280.0 131.2 L 284.9 122.9 L 289.7 115.7 L 294.6 115.7 L 299.5 115.7 L 304.4 115.7 L 309.3 115.7 L 314.2 115.7 L 319.1 115.7 L 323.9 115.7 L 328.8 115.7 L 333.7 115.7 L 338.6 115.7 L 343.5 115.7 L 348.4 115.7 L 353.3 115.7 L 356.0 115.7" fill="none" stroke="#3fb950" stroke-width="2.4"/><circle cx="289.1" cy="115.7" r="3.3" fill="#3fb950"/><path d="M 112.4 388.0 L 117.3 379.7 L 122.2 371.4 L 127.1 363.2 L 132.0 354.9 L 136.9 346.6 L 141.8 338.3 L 146.6 330.0 L 151.5 321.7 L 156.4 313.5 L 161.3 305.2 L 166.2 296.9 L 171.1 288.6 L 176.0 280.3 L 180.8 272.0 L 185.7 263.8 L 190.6 255.5 L 195.5 247.2 L 200.4 238.9 L 205.3 230.6 L 210.2 222.3 L 215.1 214.1 L 219.9 205.8 L 224.8 197.5 L 229.7 189.2 L 234.6 180.9 L 239.5 172.6 L 244.4 164.4 L 249.3 156.1 L 254.1 147.8 L 259.0 139.5 L 263.9 131.2 L 268.8 122.9 L 273.7 114.7 L 278.6 106.4 L 283.5 99.2 L 288.4 99.2 L 293.2 99.2 L 298.1 99.2 L 303.0 99.2 L 307.9 99.2 L 312.8 99.2 L 317.7 99.2 L 322.6 99.2 L 327.5 99.2 L 332.3 99.2 L 337.2 99.2 L 342.1 99.2 L 347.0 99.2 L 351.9 99.2 L 356.0 99.2" fill="none" stroke="#f85149" stroke-width="2.4"/><circle cx="282.8" cy="99.2" r="3.3" fill="#f85149"/><text x="117.8" y="372.9" font-size="11" font-weight="bold" fill="#f85149">MI300X</text><text x="133.0" y="394.3" font-size="11" font-weight="bold" fill="#3fb950">H100</text><text x="238.8" y="93.2" font-size="10" fill="#f85149">1307T</text><text x="260.4" y="129.7" font-size="10" fill="#3fb950">989T</text><rect x="452" y="74" width="278" height="314" fill="#161b22" stroke="#30363d"/><text x="591.0" y="68" text-anchor="middle" font-size="13" font-weight="bold" fill="#c9d1d9">Realized decode edge vs kernel coverage</text><line x1="452.0" y1="74" x2="452.0" y2="388" stroke="#21262d"/><text x="452.0" y="404" text-anchor="middle" font-size="11" fill="#8b98a5">80%</text><line x1="521.5" y1="74" x2="521.5" y2="388" stroke="#21262d"/><text x="521.5" y="404" text-anchor="middle" font-size="11" fill="#8b98a5">85%</text><line x1="591.0" y1="74" x2="591.0" y2="388" stroke="#21262d"/><text x="591.0" y="404" text-anchor="middle" font-size="11" fill="#8b98a5">90%</text><line x1="660.5" y1="74" x2="660.5" y2="388" stroke="#21262d"/><text x="660.5" y="404" text-anchor="middle" font-size="11" fill="#8b98a5">95%</text><line x1="730.0" y1="74" x2="730.0" y2="388" stroke="#21262d"/><text x="730.0" y="404" text-anchor="middle" font-size="11" fill="#8b98a5">100%</text><line x1="452" y1="388.0" x2="730" y2="388.0" stroke="#21262d"/><text x="446" y="392.0" text-anchor="end" font-size="11" fill="#8b98a5">0.6x</text><line x1="452" y1="330.9" x2="730" y2="330.9" stroke="#21262d"/><text x="446" y="334.9" text-anchor="end" font-size="11" fill="#8b98a5">0.8x</text><line x1="452" y1="273.8" x2="730" y2="273.8" stroke="#21262d"/><text x="446" y="277.8" text-anchor="end" font-size="11" fill="#8b98a5">1.0x</text><line x1="452" y1="216.7" x2="730" y2="216.7" stroke="#21262d"/><text x="446" y="220.7" text-anchor="end" font-size="11" fill="#8b98a5">1.2x</text><line x1="452" y1="159.6" x2="730" y2="159.6" stroke="#21262d"/><text x="446" y="163.6" text-anchor="end" font-size="11" fill="#8b98a5">1.4x</text><line x1="452" y1="102.5" x2="730" y2="102.5" stroke="#21262d"/><text x="446" y="106.5" text-anchor="end" font-size="11" fill="#8b98a5">1.6x</text><text x="591.0" y="422" text-anchor="middle" font-size="12" fill="#c9d1d9">fraction of runtime on a tuned kernel</text><line x1="452" y1="107.7" x2="730" y2="107.7" stroke="#f85149" stroke-dasharray="5,4" stroke-width="1.6"/><text x="458" y="102.7" font-size="11" fill="#f85149">hardware ceiling 1.58x</text><line x1="452" y1="273.8" x2="730" y2="273.8" stroke="#8b98a5" stroke-dasharray="2,3" stroke-width="1.3"/><text x="724" y="268.8" text-anchor="end" font-size="11" fill="#8b98a5">parity with H100</text><path d="M 452.0 371.1 L 457.6 368.9 L 463.1 366.6 L 468.7 364.3 L 474.2 361.9 L 479.8 359.5 L 485.4 356.9 L 490.9 354.4 L 496.5 351.7 L 502.0 349.0 L 507.6 346.3 L 513.2 343.4 L 518.7 340.5 L 524.3 337.5 L 529.8 334.4 L 535.4 331.2 L 541.0 327.9 L 546.5 324.6 L 552.1 321.1 L 557.6 317.5 L 563.2 313.8 L 568.8 310.1 L 574.3 306.1 L 579.9 302.1 L 585.4 297.9 L 591.0 293.6 L 596.6 289.2 L 602.1 284.6 L 607.7 279.8 L 613.2 274.9 L 618.8 269.8 L 624.4 264.5 L 629.9 259.0 L 635.5 253.3 L 641.0 247.4 L 646.6 241.3 L 652.2 234.9 L 657.7 228.2 L 663.3 221.3 L 668.8 214.0 L 674.4 206.5 L 680.0 198.6 L 685.5 190.3 L 691.1 181.7 L 696.6 172.6 L 702.2 163.1 L 707.8 153.2 L 713.3 142.7 L 718.9 131.6 L 724.4 120.0 L 730.0 107.7" fill="none" stroke="#a371f7" stroke-width="2.6"/><line x1="614.4" y1="74" x2="614.4" y2="388" stroke="#a371f7" stroke-dasharray="3,3"/><circle cx="614.4" cy="273.8" r="4" fill="#a371f7"/><text x="614.4" y="380.0" text-anchor="middle" font-size="11" font-weight="bold" fill="#a371f7">91.7%</text><text x="610.4" y="171.1" text-anchor="end" font-size="10" fill="#a371f7">below here,</text><text x="610.4" y="183.1" text-anchor="end" font-size="10" fill="#a371f7">CUDA wins</text><text x="380" y="462" text-anchor="middle" font-size="11" fill="#6e7681">Left: min(peak FLOPS, intensity x bandwidth). Right: 1.58x / [c + (1-c)x8], fallback path 8x slower.</text></svg>

The roofline shows exactly where AMD's memory advantage is not a rounding error. A kernel's ceiling is $P_\text{attain}=\min(P_\text{peak},\, I\cdot B)$, the smaller of the compute roof and the bandwidth roof, where $I$ is arithmetic intensity in FLOP per byte and $B$ is memory bandwidth. The two roofs cross at the ridge intensity $I^\star=P_\text{peak}/B$: 295 FLOP/byte for the H100, 247 for the MI300X. Autoregressive LLM decode at batch one reads every weight to produce a single token, so its intensity is barely 1 to 2 FLOP/byte, far to the left of both ridges, which means decode is bandwidth-bound and throughput scales with bandwidth alone. That hands AMD a hardware ceiling of 5.3 over 3.35, a 1.58x edge in tokens per second, before a line of model code is written. The capacity gap compounds it: a 70B model in BF16 is 140 GB of weights, which fits on a single 192 GB MI300X but needs two 80 GB H100s, so the AMD part serves memory-bound inference at lower sharding cost as well as higher bandwidth.

Prefill and training live on the far side of the ridge, where the advantage shrinks and changes character. A long-sequence attention or MLP GEMM runs at an intensity of several hundred FLOP/byte, past both ridge points, so it is compute-bound and the relevant ratio collapses to peak FLOP/s: 1,307 over 989, a 1.32x edge. But compute-bound throughput is only attainable if the software actually reaches peak, and no GEMM library reaches 100 percent. Where decode was a bandwidth number that AMD wins by physics, prefill is a utilization number that AMD wins only if rocBLAS and hipBLASLt schedule the matrix cores as tightly as cuBLAS and CUTLASS do on the NVIDIA part.

The realized advantage is a utilization multiplier, and utilization is a software quantity. Write the delivered edge as the hardware ceiling times the ratio of achieved utilizations, 1.58x scaled by $(u_\text{AMD}/u_\text{NV})$ for decode. If AMD's stack extracts 95 percent of what NVIDIA's does, the decode edge is still 1.50x; at 75 percent it is 1.19x; at the parity point $u_\text{AMD}/u_\text{NV}=0.632$ the 1.58x hardware lead is exactly cancelled and the two accelerators tie. Everything below that ratio hands the memory-bound workload back to the H100 despite AMD owning every spec sheet. The silicon sets the ceiling; the software decides how much of it arrives.

The specification table reads as a clean sweep, which is precisely the trap. Every row below favors the MI300X, and every row describes a ceiling that only a mature software stack can reach.

SpecificationAMD MI300XNVIDIA H100 SXMAMD edge
HBM3 capacity192 GB80 GB2.40x
Memory bandwidth5.3 TB/s3.35 TB/s1.58x
Dense BF16 throughput1,307 TFLOP/s989 TFLOP/s1.32x
Ridge intensity247 FLOP/byte295 FLOP/bytelower is better
GPUs to hold a 70B model12half the parts

Coverage, not tuned-kernel speed, is where the software surface area actually bites. Model a workload as a fraction $c$ of runtime that hits a well-tuned kernel at hardware parity and a residual $(1-c)$ that falls to an unfused or reference path running $s$ times slower; the total time relative to a fully covered stack is $T_\text{rel}=c+(1-c)\,s$. With a fallback only 8x slower, 99 percent coverage costs 7 percent, 95 percent coverage costs 35 percent, and 90 percent coverage costs 70 percent. Setting that penalty equal to the 1.58x hardware edge, the break-even coverage is 91.7 percent: above it AMD's memory advantage survives the software tax, below it CUDA wins the same benchmark on slower silicon. This is why a stack can match NVIDIA on the ten kernels a MLPerf run exercises and still lose a real training job that touches four hundred, and why the single most valuable ROCm number is not a FLOP count but the percentage of a framework's operators that have a tuned kernel.

The coverage table converts the software gap into the currency that matters, delivered throughput. It reads the erasure curve on the right panel above: how much of the 1.58x decode ceiling actually reaches the user as a function of how complete the tuned-kernel coverage is.

Tuned-kernel coverageSlowdown vs full coverageRealized decode edge
100%1.00x1.58x
97%1.21x1.31x
95%1.35x1.17x
91.7%1.58x1.00x (tie)
90%1.70x0.93x
85%2.05x0.77x

HIP makes the port look free and then charges for the last mile. HIP is a near line-for-line clone of the CUDA runtime API, so the hipify tool mechanically rewrites roughly 95 percent of call sites; a 4,200-call kernel repository converts to about 3,990 automatically and leaves 210 by hand. The residual is where the cost concentrates, and its most common cause is that CDNA hardware runs a 64-lane wavefront where NVIDIA runs a 32-thread warp, so any kernel that hard-codes a warp width of 32 in a shuffle, a ballot or a shared-memory tiling assumption is silently half-occupied or wrong until it is rewritten. The API port is a weekend; the performance port is a quarter.

The libraries are the moat, not the instruction set. NVIDIA's advantage is the decade of tuning frozen into cuBLAS, cuDNN, CUTLASS and the fused-attention kernels that PyTorch, vLLM and Triton call by default, and AMD's task is to reach coverage parity in rocBLAS, MIOpen and Composable Kernel across the same operator surface. The trend is real: PyTorch ships upstream ROCm wheels, Triton emits AMD GCN, vLLM lists MI300X as a first-class target, and the two largest supercomputers ever built, Frontier at Oak Ridge and El Capitan at Livermore, run more than 37,000 and 43,000 Instinct GPUs respectively on this exact stack. ROCm at exascale is a proof that the coverage gap is an engineering backlog, not a hardware verdict.

The two equations that govern the whole comparison are compact enough to keep in view at once, the roofline and the Amdahl coverage penalty:

$$P_\text{attain}=\min\!\left(P_\text{peak},\; I\cdot B\right), \qquad T_\text{rel}=c+(1-c)\,s.$$

The first says AMD's bandwidth wins wherever intensity is low; the second says a thin band of uncovered operators, amplified by a slow fallback, is enough to overturn it.

Read ROCm through a software surface area lens rather than a FLOP/s lens, and the market stops being a paradox. The MI300X is the faster memory system and the roomier one, so on the bandwidth-bound half of modern AI it starts 1.58x ahead and can hold models an H100 must split. What it does not yet start with is a decade of tuned kernels covering every corner of the operator space, and the roofline is unforgiving about that: at 90 percent coverage a 1.70x software penalty erases a 1.58x hardware lead and the slower silicon wins the benchmark. Every hard problem in this stack is the same problem wearing a different mask, whether it is a missing fused-attention kernel, a warp width of 32 baked into someone else's CUDA, a collective that has not been tuned for the Infinity Fabric mesh, or a framework that defaults to a cuDNN path with no ROCm equal. None of them are deficits in the transistors. They are all coverage of a moving software target, which is why ROCm's progress is measured not in TFLOP/s but in the shrinking fraction of a real workload that still falls off the fast path.

rocmamd instinctamd gpu computeradeon computerocblasmiopencomposable kernelhipifycdna architecturemi300x rocm

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.