Llama 2 70B · Offline · v4.1 → v5.0 → v5.1 → v6.0 · same GPU, tracked across every MLPerf round
Nobody buys new silicon between MLPerf rounds — the H100 in a datacenter today is the same H100 that
was there a year ago. But the tokens/sec it delivers keeps climbing anyway, purely from better kernels,
quantization, and inference-serving software. That's throughput procurement already paid for and hasn't
claimed. This page tracks it directly: the same three flagship GPUs, the same workload, every MLPerf
round since we started tracking. The first chart is raw speed. The second holds today's price constant
and asks what a token actually costs — which is where the story gets less predictable.
+12.7%
B200-SXM tok/s·GPU, v4.1→v6.0 — same silicon, pure software/firmware gain
4.1×
more tok/s per GPU — earliest H100 result vs latest B200
9–22%
range of decline in $/M tokens across flagships since v4.1, at today's floor price
2
MLPerf rounds since H100-SXM last got a llama2-70b submission (v5.0)
Throughput — tokens/sec per GPU, by MLPerf round
H100 SXMH200 SXMB200 SXM
Cost per million tokens — at today’s floor price
H100 SXMH200 SXMB200 SXM
B200 SXM costs $0.146 per million tokens at today's floor
pricing — more than H200 SXM at $0.144 — despite B200 SXM being
the newer, faster chip (12,699 vs 4,415 tok/s·GPU). B200 SXM's price premium
over H200 SXM outpaces the throughput it delivers.
GB200 NVL72's per-GPU throughput (12,334 tok/s, v6.0) is
in line with, not ahead of standalone B200-SXM (12,699 tok/s, v6.0) —
the rack's real advantage is 72 GPUs sharing one NVLink domain, not faster silicon per chip.
No pricing data yet for GB200, so it's excluded from the charts above.