NVIDIA GB200
sm_100
Records held
86
Runs
93
Operations
55 implementations
Records held on this GPU
Operation · workload
History
Record
Implementation
Since
Show all 40 records ›Showing all 40 records ⌄
GQA paged prefill causal h5 kv1 d128 ps1bf16 · [1, 5, 128] · num pages 33417 · num kv indices 19
170.8µs
2026-04-12
GQA paged prefill causal h5 kv1 d128 ps1bf16 · [1, 5, 128] · num pages 167 · num kv indices 33
105.3µs
2026-04-12
46 more records · Open in the records ledger →
Coverage by operation family
Family
Operations
Runs
With source
gqa-paged-attention
4
73
73
gqa-ragged-attention
1
20
20
Serving evidence
End-to-end serving results measured on this hardware — a separate corpus with its own resolver, never ranked against the kernel records above.
Model · workload
Scenario
Best reported
Runs
Llama-2-70B (MLPerf reference)llama2-70b-99 · Server
server
49,360 tok/s
3
Llama-2-70B (MLPerf reference)llama2-70b-99 · Offline
offline
52,062 tok/s
3
Llama-3.1-405B (MLPerf reference)llama3.1-405b · Offline
offline
856 tok/s
2
DeepSeek-R1 (MLPerf reference)deepseek-r1 · Server
server
336,106 tok/s
2
Llama-3.1-405B (MLPerf reference)llama3.1-405b · Server
server
596 tok/s
2
Llama-2-70B (MLPerf reference)llama2-70b-99 · Server
server
49,360 tok/s
2
DeepSeek-R1 (MLPerf reference)deepseek-r1 · Offline
offline
486,141 tok/s
2
Llama-2-70B (MLPerf reference)llama2-70b-99 · Offline
offline
51,737 tok/s
2
42 serving runs across 32 comparison groups · Open in the serving resolver →
Record activity
Nov
Dec
Jan
Feb
Mar
91Apr
May
Jun
Jul
Aug
Sep
Oct
Sources: FlashInfer-Bench (2026-04-12) · Apache-2.0last observed 2026-04-12