The fastest known GPU kernel for your exact workload.
Same shapes, same dtype, same GPU, same protocol. Anything else is not ranked against it.
Batched compact-Householder QRBatched symmetric eigendecompositiondeepseek-r1Query syntax| Operation / workload | Implementation | Latency | Hardware |
|---|---|---|---|
| Batched compact-Householder QRsuite of 12 cases | submission 826595 | 422.4µs 73.8% faster · was 1.62 ms | B200 |
| Batched symmetric eigendecompositionsuite of 13 cases | submission 839675 | 0ns 100.0% faster · was 47.7 ms | B200 |
| Batched Cholesky factorizationsuite of 15 cases | submission 928092 | 2.61µs 87.7% faster · was 21.2 µs | B200 |
| Mamba2ReturnYfp32 · [2048, 128, 8, 64] | torch.compile (inductor) | 7.47ms 39.3% faster · was 12.3 ms | H100 |
| Mamba2ReturnFinalStatefp32 · [2048, 128, 8, 64] | torch.compile (inductor) | 2.66ms 71.0% faster · was 9.18 ms | H100 |
| ReLUSelfAttentionfp32 · [16, 1024, 768] | torch.compile (inductor) | 4.97ms 35.0% faster · was 7.65 ms | H100 |
| MinGPTCausalAttentionfp32 · [128, 512, 768] | torch.compile (inductor) | 15.7ms 20.3% faster · was 19.7 ms | H100 |
| RMSNorm h512bf16 · [512] · batch 7 | gpt-o3 / triton19c647 | 6.19µs 1.7% faster · was 6.29 µs | B200 |
Shown as published by their sources. Evidence levels ↓
Where the numbers come from
GPU MODE KernelBot9,585 runs · 2026-09-28
NVIDIA SOL-ExecBench8,024 runs · 2026-08-24 · stale
FlashInfer-Bench5,267 runs · 2026-09-07 · stale
Liger-Kernel benchmarks1,649 runs · 2026-08-18 · stale
kernelbench988 runs · 2026-08-26
Evidence levels
Reproduction-ready 425Reported 25,088
Every result is imported from its source and shown as published. KernelIndex has not independently rerun any of them, so no result is Verified yet. Sources and limitations →