Skip to content
KernelIndex
Search⌘K

GEMM n7168 k5120

86 eligible runs
gemm

General matrix multiply (GEMM) C = A @ B.T. Captured from Qwen3 14B qkv_proj (combined Q+K+V, (40+8+8)*128=7168, hidden=5120).

Source baseline · unbeaten
18.1µsmean
flashinfer / wrapper4c2606FlashInfer-Bench baselines · Apache-2.0 · python

Reported evidence · last observed 2026-03-24. The source's designated baseline implementation. Reported by source; not independently reproduced.

Current records

Source-native comparison · GPU NVIDIA B200 · Workload m = 7 · fp16 · CUDA 12.8 · Framework pytorch 2.9.1+cu128 · Protocol flashinfer-bench · mean · 2 results · last observed 2026-03-24Record history →
Estimated floor 9.18 µs · record 1.97× above itestimate, not evidence ›
DRAM 9.18 µs · compute 228 ns · bandwidth-bound on B200
every declared tensor crosses HBM exactly once (8,000 GB/s, B200 datasheet)
2·M·N·K with M=7, N=7168, K=5120 at the dense bf16 peak (2,250 TFLOP/s)
headroom-v1: a lower bound from declared tensors and datasheet peaks. A kernel can sit well above it for good reasons.
#
Implementation
Latency
vs #1
Trust
Observed
1
flashinfer / wrapper4c2606baselineFlashInfer-Bench baselines
18.1µs
1.00×
Reported · Apache-2.0 · source
2026-03-24stale

Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.

source mirroredApache-2.0no install recipeView source →Run detail →
2
flashinfer / wrapper4c2606baselineFlashInfer-Bench baselines
18.4µs
1.01×
Reported · Apache-2.0 · source
2026-03-23stale

Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.

source mirroredApache-2.0no install recipeView source →Run detail →

Scaling by m

100.0 µs14248825620538km →19.1 µs19.2 µs18.4 µs18.1 µs18.4 µs19.2 µs19.4 µs18.5 µs20.0 µs20.2 µs20.2 µs20.4 µs20.2 µs20.3 µs21.1 µs21.0 µs20.8 µs20.4 µs20.3 µs19.7 µs19.6 µs19.6 µs19.6 µs23.6 µs23.7 µs21.5 µs21.4 µs25.9 µs25.8 µs22.7 µs22.6 µs26.1 µs26.2 µs23.4 µs23.1 µs23.4 µs23.3 µs23.2 µs23.3 µs50.4 µs97.2 µs110.4 µs405.8 µsflashinfer / wrapper 4c2606
flashinfer / wrapper 4c2606best per workload · NVIDIA B200 · CUDA 12.8 · pytorch 2.9.1+cu128 · flashinfer-bench held constant · log scale

Implementations

Implementation
Runtime
Best latency
Evidence
Availability
flashinfer / wrapper4c2606FlashInfer-Bench baselines
python
18.1µs
1.00×
Reported
Apache-2.0 · source

Semantics

Inputs and outputs
afp16 [m, k]
bfp16 [n, k]
cfp16 [m, n]
Axes and behavior
kconstant = 5120
mvariable
nconstant = 7168
determinismunspecified
constraintsNo mutation or aliasing
Identity
aliasgemm_n7168_k5120modelqwen3-14bsha2561a01e1c673a4…
Sources: FlashInfer-Bench (2026-03-24) · Apache-2.0last observed 2026-03-24How records are decidedJSON