Skip to content
KernelIndex
Search
Models
Records
GPUs
Docs
More
Records
GPUs
Docs
Sign in
Kernels
Serving
⏎
Kernels
Serving
⏎
dtype bf16
✕
20 operations match
Pick one to compare its implementations; your filters carry over
Operation
Family
Published runs
Last run
FP8 MoE expert linear
none on bf16
moe
33
2026-08-23
Grouped topk MoE routing backward
none on bf16
moe
32
2026-08-23
MoE expert batched execution with capacity factor
none on bf16
moe
24
2026-08-19
MoE expert computation
none on bf16
moe
27
2026-08-22
MoE expert computation with weighted accumulation
none on bf16
moe
23
2026-08-20
MoE expert inference batched dispatch
none on bf16
moe
25
2026-08-21
MoE expert load balancing and token capacity backward
none on bf16
moe
22
2026-07-23
MoE expert MLP with load balancing
none on bf16
moe
28
2026-08-21
MoE expert parallel execution
none on bf16
moe
24
2026-08-21
MoE expert parallel execution backward
none on bf16
moe
34
2026-08-21
MoE expert parallel execution with weighted aggregation
none on bf16
moe
22
2026-08-20
MoE expert SwiGLU with down projection backward
none on bf16
moe
48
2026-08-24
MoE expert token radix sort with prefix sum
none on bf16
moe
30
2026-08-24
MoE group score aggregation and masking
none on bf16
moe
19
2026-07-23
MoE layer complete forward with residual
none on bf16
moe
30
2026-08-23
MoE sparse routing and dispatch
none on bf16
moe
30
2026-08-16
MoE sparse routing and dispatch
none on bf16
moe
31
2026-08-20
MoE sparse routing and dispatch backward
none on bf16
moe
22
2026-08-23
MoE training token repeat and expert computation
none on bf16
moe
23
2026-08-23
Sparse MoE routing and experts
none on bf16
moe
22
2026-08-18
Ranked only against runs that measured the same thing.