Skip to content
KernelIndex
Search⌘K

GEMM n5120 k5120

86 eligible runs
gemm

General matrix multiply (GEMM) C = A @ B.T. Captured from Qwen3 14B o_proj (q_heads*head_dim=40*128=5120 → hidden=5120). Square GEMM.

Source baseline · unbeaten
14.5µsmean
flashinfer / wrapperad9a00FlashInfer-Bench baselines · Apache-2.0 · python

Reported evidence · last observed 2026-03-24. The source's designated baseline implementation. Reported by source; not independently reproduced.

Current records

Source-native comparison · GPU NVIDIA B200 · Workload m = 7 · fp16 · CUDA 12.8 · Framework pytorch 2.9.1+cu128 · Protocol flashinfer-bench · mean · 2 results · last observed 2026-03-24Record history →
Estimated floor 6.56 µs · record 2.21× above itestimate, not evidence ›
DRAM 6.56 µs · compute 163 ns · bandwidth-bound on B200
every declared tensor crosses HBM exactly once (8,000 GB/s, B200 datasheet)
2·M·N·K with M=7, N=5120, K=5120 at the dense bf16 peak (2,250 TFLOP/s)
headroom-v1: a lower bound from declared tensors and datasheet peaks. A kernel can sit well above it for good reasons.
#
Implementation
Latency
vs #1
Trust
Observed
1
flashinfer / wrapperad9a00baselineFlashInfer-Bench baselines
14.5µs
1.00×
Reported · Apache-2.0 · source
2026-03-24stale

Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.

source mirroredApache-2.0no install recipeView source →Run detail →
2
flashinfer / wrapperad9a00baselineFlashInfer-Bench baselines
14.6µs
1.00×
Reported · Apache-2.0 · source
2026-03-23stale

Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.

source mirroredApache-2.0no install recipeView source →Run detail →

Scaling by m

100.0 µs200.0 µs14248825620538km →14.7 µs14.6 µs14.5 µs14.5 µs14.7 µs16.2 µs16.4 µs15.6 µs15.5 µs15.0 µs14.9 µs14.9 µs14.7 µs14.9 µs16.5 µs16.4 µs16.9 µs16.4 µs16.3 µs15.5 µs15.3 µs15.4 µs15.5 µs21.8 µs22.0 µs23.1 µs23.4 µs22.2 µs23.2 µs23.2 µs23.5 µs20.1 µs19.3 µs19.1 µs19.0 µs18.9 µs18.8 µs18.9 µs18.8 µs37.8 µs79.9 µs98.7 µs281.0 µsflashinfer / wrapper ad9a00
flashinfer / wrapper ad9a00best per workload · NVIDIA B200 · CUDA 12.8 · pytorch 2.9.1+cu128 · flashinfer-bench held constant

Implementations

Implementation
Runtime
Best latency
Evidence
Availability
flashinfer / wrapperad9a00FlashInfer-Bench baselines
python
14.5µs
1.00×
Reported
Apache-2.0 · source

Semantics

Inputs and outputs
afp16 [m, k]
bfp16 [n, k]
cfp16 [m, n]
Axes and behavior
kconstant = 5120
mvariable
nconstant = 5120
determinismunspecified
constraintsNo mutation or aliasing
Identity
aliasgemm_n5120_k5120modelqwen3-14bsha256f07beba5e07f…
Sources: FlashInfer-Bench (2026-03-24) · Apache-2.0last observed 2026-03-24How records are decidedJSON