NVIDIA H100
sm_90
Records held
823
Runs
2,051
Operations
271848 implementations
Records held on this GPU
Operation · workload
History
Record
Implementation
Since
Show all 40 records ›Showing all 40 records ⌄
783 more records · Open in the records ledger →
Coverage by operation family
Family
Operations
Runs
With source
conv
97
419
417
gemm
56
248
243
model
50
200
200
neighborhood-attention
1
168
168
rmsnorm
3
103
103
moe
1
84
84
elementwise
2
80
71
reduction
10
73
68
swiglu
1
72
72
trimul
1
71
70
cross-entropy
2
58
58
jsd
3
56
56
activation
13
52
52
relu-squared
1
48
48
rope
1
48
48
scan
5
39
39
tvd
1
36
36
layernorm
2
34
34
kl-divergence
2
28
28
sort
1
26
21
histogram
1
24
22
kto
1
20
20
loss
4
16
16
normalization
3
12
12
pooling
3
12
12
softmax
2
8
8
groupnorm
1
4
4
batchnorm
1
4
4
attention
1
4
4
instancenorm
1
4
4
Serving evidence
End-to-end serving results measured on this hardware — a separate corpus with its own resolver, never ranked against the kernel records above.
Model · workload
Scenario
Best reported
Runs
Llama-2-70B (MLPerf reference)llama2-70b-99 · Offline
offline
62,352 tok/s
1
Llama-2-70B (MLPerf reference)llama2-70b-99 · Offline
offline
124,879 tok/s
1
Llama-3.1-405B (MLPerf reference)llama3.1-405b · Server
server
1,127 tok/s
1
Llama-2-70B (MLPerf reference)llama2-70b-99 · Offline
offline
62,352 tok/s
1
Llama-2-70B (MLPerf reference)llama2-70b-99 · Server
server
122,274 tok/s
1
Llama-2-70B (MLPerf reference)llama2-70b-99 · Server
server
61,101 tok/s
1
Llama-3.1-405B (MLPerf reference)llama3.1-405b · Offline
offline
783 tok/s
1
Llama-3.1-8B (MLPerf reference)llama3.1-8b · Server
server
5,104 tok/s
1
14 serving runs across 14 comparison groups · Open in the serving resolver →
Record activity
25Nov
6Dec
3Jan
6Feb
857Mar
61Apr
May
43Jun
31Jul
Aug
Sep
Oct
Sources: Liger-Kernel benchmarks (2026-07-26) · BSD-2-Clause · KernelBench baseline timings (2026-03-05) · MIT · GPU MODE KernelBot (2026-05-08) · June 9 Researcher Reciprocity License v1.0last observed 2026-07-26