Skip to content
KernelIndex
Search⌘K

The fastest known GPU kernel for your exact workload.

Same shapes, same dtype, same GPU, same protocol. Anything else is not ranked against it.

Batched compact-Householder QRBatched symmetric eigendecompositiondeepseek-r1Query syntax

Liger2.10ms15.02× faster than baselineUse it →

Reported evidence · observed 2026-04-02 · BSD-2-Clause

pip install "liger-kernel==0.7.0"
Operation / workloadImplementationLatencyHardware
Batched compact-Householder QRsuite of 12 casessubmission 826595422.4µs
73.8% faster · was 1.62 ms
B200
Batched symmetric eigendecompositionsuite of 13 casessubmission 8396750ns
100.0% faster · was 47.7 ms
B200
Batched Cholesky factorizationsuite of 15 casessubmission 9280922.61µs
87.7% faster · was 21.2 µs
B200
Mamba2ReturnYfp32 · [2048, 128, 8, 64]torch.compile (inductor)7.47ms
39.3% faster · was 12.3 ms
H100
Mamba2ReturnFinalStatefp32 · [2048, 128, 8, 64]torch.compile (inductor)2.66ms
71.0% faster · was 9.18 ms
H100
ReLUSelfAttentionfp32 · [16, 1024, 768]torch.compile (inductor)4.97ms
35.0% faster · was 7.65 ms
H100
MinGPTCausalAttentionfp32 · [128, 512, 768]torch.compile (inductor)15.7ms
20.3% faster · was 19.7 ms
H100
RMSNorm h512bf16 · [512] · batch 7gpt-o3 / triton19c6476.19µs
1.7% faster · was 6.29 µs
B200

Shown as published by their sources. Evidence levels ↓

Where the numbers come from
GPU MODE KernelBot9,585 runs · 2026-09-28
NVIDIA SOL-ExecBench8,024 runs · 2026-08-24 · stale
FlashInfer-Bench5,267 runs · 2026-09-07 · stale
Liger-Kernel benchmarks1,649 runs · 2026-08-18 · stale
kernelbench988 runs · 2026-08-26
Evidence levels
Reproduction-ready 425Reported 25,088

Every result is imported from its source and shown as published. KernelIndex has not independently rerun any of them, so no result is Verified yet. Sources and limitations →