GEMM NVFP4 n4096 k4096
NVFP4 dense GEMM C = (A @ B.T) * alpha (N=4096, K=4096). Inputs are NVFP4-quantized: packed E2M1 data (2 values/byte, int8 storage) plus per-16 UE4M3 block scales (swizzled 128x4 layout, int8 storage), and a scalar global de-scale alpha = 1/(global_sf_a*global_sf_b). The reference dequantizes both operands and matmuls, so NVFP4 kernels are numerically exact against it.
Reported evidence · last observed 2026-06-06. The source's designated baseline implementation. Reported by source; not independently reproduced.
Current records
Not measured on H100 for this workload. Challenges →
Estimated floor 1.25 µs · record 77.06× above itestimate, not evidence ›
196.6µs1.00×Reported · Apache-2.0 · source2026-06-06
Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.
297.8µs1.01×Reported · Apache-2.0 · source2026-06-06
Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.
3100.3µs1.04×Reported · Apache-2.0 · source2026-06-06
Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.
4103.0µs1.07×Reported · Apache-2.0 · source2026-06-06
Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.
5129.4µs1.34×Reported · Apache-2.0 · source2026-06-06
Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.
6135.9µs1.41×Reported · Apache-2.0 · source2026-06-06
Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.