llama-3-1-8b
0 of 14 operations have a usable best known on NVIDIA B200 · 14 without one
Best known on NVIDIA B200
GEMM n28672 k4096fp16 · [8, 4096]torch_matmul_655587not usable55.4µsReported · Apache-2.0 · source2025-10-16stale
No entry in this cohort passes the usability policy (no install recipe); the fastest known is shown. Measured by the source's own harness. Ranked by its metric under ranking-v1. 7 more entries in this cohort.
GEMM n4096 k14336fp16 · [16, 14336]torch_matmul_254647not usable35.6µsReported · Apache-2.0 · source2025-10-20stale
No entry in this cohort passes the usability policy (no install recipe); the fastest known is shown. Measured by the source's own harness. Ranked by its metric under ranking-v1. 5 more entries in this cohort.
GEMM n4096 k4096fp16 · [128, 4096]torch / matmul0d13dfnot usable16.1µsReported · Apache-2.0 · source2025-10-16stale
No entry in this cohort passes the usability policy (no install recipe); the fastest known is shown. Measured by the source's own harness. Ranked by its metric under ranking-v1. 8 more entries in this cohort.
GEMM n6144 k4096fp16 · [7, 4096]gemini-2.5-pro / cuda4bc599not usable18.3µsReported · Apache-2.0 · source2025-10-16stale
No entry in this cohort passes the usability policy (no install recipe); the fastest known is shown. Measured by the source's own harness. Ranked by its metric under ranking-v1. 5 more entries in this cohort.
GQA paged decode h32 kv8 d128 ps1bf16 · [1, 32, 128] · num pages 18 · num kv indices 2gpt-5 / cuda95c7fenot usable8.19µsReproduction-ready · Apache-2.0 · source2025-10-16stale
No entry in this cohort passes the usability policy (no install recipe); the fastest known is shown. Measured by the source's own harness. Ranked by its metric under ranking-v1. 6 more entries in this cohort.
GQA paged decode h32 kv8 d128 ps64bf16 · [4, 32, 128] · num pages 349 · num kv indices 212flashinfer / wrapperad4135not usable28.8µsReported · Apache-2.0 · source2026-04-30
No entry in this cohort passes the usability policy (no install recipe); the fastest known is shown. Measured by the source's own harness. Ranked by its metric under ranking-v1.
GQA paged prefill causal h32 kv8 d128 ps1bf16 · [2, 32, 128]gpt-o3 / cudad4241dnot usable31.5µsReported · Apache-2.0 · source2025-10-21stale
No entry in this cohort passes the usability policy (no install recipe); the fastest known is shown. Measured by the source's own harness. Ranked by its metric under ranking-v1. 6 more entries in this cohort.
Fused add RMSNorm h4096bf16 · [4096] · batch 7gemini-2.5-pro / tritondc28mjnot usable6.64µsReported · Apache-2.0 · source2025-10-16stale
No entry in this cohort passes the usability policy (no install recipe); the fastest known is shown. Measured by the source's own harness. Ranked by its metric under ranking-v1. 7 more entries in this cohort.
RMSNorm h4096bf16 · [4096] · batch 16gemini-2.5-pro / cudaaaf481not usable7.27µsReproduction-ready · Apache-2.0 · source2025-10-16stale
No entry in this cohort passes the usability policy (no install recipe); the fastest known is shown. Measured by the source's own harness. Ranked by its metric under ranking-v1. 5 more entries in this cohort.
GQA ragged prefill causal h32 kv8 d128bf16 · [35, 8, 128] · #10e83cflashinfer / wrapperf9a07bnot usable11.0µsReported · Apache-2.0 · source2025-10-21stale
No entry in this cohort passes the usability policy (no install recipe); the fastest known is shown. Measured by the source's own harness. Ranked by its metric under ranking-v1. 19 more entries in this cohort.
Coverage gaps on NVIDIA B200
No eligible evidence yet for these operations. Challenges →