Skip to content
KernelIndex
Search⌘K

NVIDIA GB200

sm_100
Records held
86
Runs
93
Operations
55 implementations

Records held on this GPU

Operation · workload
History
Record
Implementation
Since
Show all 40 records ›
GQA paged prefill causal h5 kv1 d128 ps1bf16 · [1, 5, 128] · num pages 33417 · num kv indices 19
170.8µs
2026-04-12
GQA paged prefill causal h5 kv1 d128 ps1bf16 · [1, 5, 128] · num pages 167 · num kv indices 33
105.3µs
2026-04-12

46 more records · Open in the records ledger →

Coverage by operation family

Family
Operations
Runs
With source
gqa-paged-attention
4
73
73
gqa-ragged-attention
1
20
20

Serving evidence

End-to-end serving results measured on this hardware — a separate corpus with its own resolver, never ranked against the kernel records above.

Model · workload
Scenario
Best reported
Runs
Llama-2-70B (MLPerf reference)llama2-70b-99 · Server
server
49,360 tok/s
3
Llama-2-70B (MLPerf reference)llama2-70b-99 · Offline
offline
52,062 tok/s
3
Llama-3.1-405B (MLPerf reference)llama3.1-405b · Offline
offline
856 tok/s
2
DeepSeek-R1 (MLPerf reference)deepseek-r1 · Server
server
336,106 tok/s
2
Llama-3.1-405B (MLPerf reference)llama3.1-405b · Server
server
596 tok/s
2
Llama-2-70B (MLPerf reference)llama2-70b-99 · Server
server
49,360 tok/s
2
DeepSeek-R1 (MLPerf reference)deepseek-r1 · Offline
offline
486,141 tok/s
2
Llama-2-70B (MLPerf reference)llama2-70b-99 · Offline
offline
51,737 tok/s
2

42 serving runs across 32 comparison groups · Open in the serving resolver →

Record activity

Nov
Dec
Jan
Feb
Mar
91Apr
May
Jun
Jul
Aug
Sep
Oct
Sources: FlashInfer-Bench (2026-04-12) · Apache-2.0last observed 2026-04-12