Skip to content
KernelIndex
Search⌘K

NVIDIA B200

sm_100
Records held
1,926
Runs
17,071
Operations
33712,066 implementations

Records held on this GPU

Operation · workload
History
Record
Implementation
Since
422.4µs
2026-06-22
2.61µs
2026-07-29
RMSNorm h7168suite of 8 cases
6.48µs
2026-07-02
RMSNorm h7168suite of 8 cases
6.15µs
2026-07-30
FP8 shared expert MLPsuite of 16 cases
167.2µs
2026-08-23
RMSNorm h4096suite of 14 cases
6.33µs
2026-06-07
RMSNorm h1536suite of 8 cases
3.62µs
2026-06-06
RMSNorm h2048suite of 7 cases
4.12µs
2026-06-07
RMSNorm h512suite of 8 cases
2.84µs
2026-07-02
RMSNorm h128suite of 14 cases
3.79µs
2026-06-07
Show all 40 records ›
29.5µs
2026-08-19
7.94µs
2026-07-02
379.2µs
2026-05-19
288.8µs
2026-08-21
255.6µs
2026-07-02
233.2µs
2026-08-19
14.0µs
2026-06-15
GEMM n28672 k4096suite of 43 cases
63.2µs
2026-06-14
GEMM n5120 k2048suite of 25 cases
19.2µs
2026-05-25
GEMM n5120 k2048suite of 25 cases
18.6µs
2026-08-02
GEMM n4096 k14336suite of 43 cases
39.9µs
2026-06-08
GEMM n2048 k4096suite of 29 cases
23.0µs
2026-07-03
GEMM n256 k7168suite of 17 cases
9.24µs
2026-06-07
751.0µs
2026-08-23
GEMM n6144 k4096suite of 43 cases
19.6µs
2026-06-08
GEMM n128 k2048suite of 25 cases
5.73µs
2026-06-07
Fused add RMSNorm h7168suite of 8 cases
7.78µs
2026-07-02
Fused add RMSNorm h7168suite of 8 cases
7.16µs
2026-08-15
Fused add RMSNorm h4096suite of 14 cases
7.89µs
2026-07-01
Fused add RMSNorm h4096suite of 14 cases
7.10µs
2026-07-23
Fused add RMSNorm h2048suite of 7 cases
4.81µs
2026-06-07
121.0µs
2026-06-06

1886 more records · Open in the records ledger →

Coverage by operation family

Family
Operations
Runs
With source
gemm
49
4,463
3,574
rmsnorm
49
2,133
669
gqa-paged-attention
26
1,474
1,285
moe
41
1,215
57
other
25
888
0
rope
20
690
0
gemv
1
678
385
mla-paged-attention
7
653
376
mlp
20
650
0
qr
1
515
183
attention
14
473
0
quantization
14
438
17
vision
14
372
0
cholesky
1
337
107
gqa-ragged-attention
8
322
238
ssm
11
322
0
eigh
1
286
109
audio
7
192
0
elementwise
2
150
126
decoder
4
145
0
diffusion
6
144
0
reduction
1
88
63
video
3
79
0
gated-deltanet
3
71
68
conv
2
64
58
histogram
1
54
52
trimul
1
43
41
groupnorm
1
42
42
jsd
1
24
24
sort
1
23
21
scan
1
23
21
mla-ragged
1
20
20

Serving evidence

End-to-end serving results measured on this hardware — a separate corpus with its own resolver, never ranked against the kernel records above.

Model · workload
Scenario
Best reported
Runs
Llama-3.1-405B (MLPerf reference)llama3.1-405b · Offline
offline
1,661 tok/s
13
Llama-3.1-405B (MLPerf reference)llama3.1-405b · Server
server
1,280 tok/s
13
Llama-2-70B (MLPerf reference)llama2-70b-99 · Offline
offline
102,909 tok/s
11
Llama-2-70B (MLPerf reference)llama2-70b-99 · Server
server
101,611 tok/s
11
Llama-2-70B (MLPerf reference)llama2-70b-99 · Offline
offline
102,909 tok/s
10
Llama-3.1-405B (MLPerf reference)llama3.1-405b · Interactive
interactive
771 tok/s
10
Llama-2-70B (MLPerf reference)llama2-70b-99 · Server
server
101,611 tok/s
10
Llama-3.1-8B (MLPerf reference)llama3.1-8b · Server
server
128,794 tok/s
9

151 serving runs across 31 comparison groups · Open in the serving resolver →

Record activity

65Nov
34Dec
40Jan
43Feb
585Mar
723Apr
272May
421Jun
882Jul
247Aug
Sep
Oct
Sources: Liger-Kernel benchmarks (2026-05-27) · BSD-2-Clause · FlashInfer-Bench (2026-06-06) · Apache-2.0 · NVIDIA SOL-ExecBench (2026-08-24) · GPU MODE KernelBot (2026-07-29) · June 9 Researcher Reciprocity License v1.0last observed 2026-08-24