Skip to content
KernelIndex
Search
Models
Records
GPUs
Docs
More
Records
GPUs
Docs
Sign in
Kernels
Serving
⏎
Kernels
Serving
⏎
gpu B200
✕
20 operations match
Pick one to compare its implementations; your filters carry over
Operation
Family
Published runs
Last run
FP8 MoE expert linear
33 on B200 · best 93.6 µs
moe
33
2026-08-23
Grouped topk MoE routing backward
32 on B200 · best 23.0 µs
moe
32
2026-08-23
MoE expert batched execution with capacity factor
24 on B200 · best 2.14 ms
moe
24
2026-08-19
MoE expert computation
27 on B200 · best 3.67 ms
moe
27
2026-08-22
MoE expert computation with weighted accumulation
23 on B200 · best 373.4 µs
moe
23
2026-08-20
MoE expert inference batched dispatch
25 on B200 · best 1.21 ms
moe
25
2026-08-21
MoE expert load balancing and token capacity backward
22 on B200 · best 7.21 µs
moe
22
2026-07-23
MoE expert MLP with load balancing
28 on B200 · best 713.5 µs
moe
28
2026-08-21
MoE expert parallel execution
24 on B200 · best 2.37 ms
moe
24
2026-08-21
MoE expert parallel execution backward
34 on B200 · best 12.6 ms
moe
34
2026-08-21
MoE expert parallel execution with weighted aggregation
22 on B200 · best 4.93 ms
moe
22
2026-08-20
MoE expert SwiGLU with down projection backward
48 on B200 · best 42.2 µs
moe
48
2026-08-24
MoE expert token radix sort with prefix sum
30 on B200 · best 5.13 µs
moe
30
2026-08-24
MoE group score aggregation and masking
19 on B200 · best 3.51 µs
moe
19
2026-07-23
MoE layer complete forward with residual
30 on B200 · best 1.00 ms
moe
30
2026-08-23
MoE sparse routing and dispatch
30 on B200 · best 5.74 ms
moe
30
2026-08-16
MoE sparse routing and dispatch
31 on B200 · best 218.3 µs
moe
31
2026-08-20
MoE sparse routing and dispatch backward
22 on B200 · best 15.2 µs
moe
22
2026-08-23
MoE training token repeat and expert computation
23 on B200 · best 4.47 ms
moe
23
2026-08-23
Sparse MoE routing and experts
22 on B200 · best 740.0 µs
moe
22
2026-08-18
Ranked only against runs that measured the same thing.