Grouped GEMM NVFP4 m contiguous g4 n2048 k2048
M-contiguous grouped NVFP4 GEMM (4 groups, N=2048, K=2048). Rows of all groups concatenated; m_indptr gives group boundaries. NVFP4: packed E2M1 (int8) + UE4M3 per-16 block scale (logical layout, int8) + per-group global de-scale alpha. Hand-written CUDA baseline (runs on sm100, unlike the sm120-only flashinfer grouped kernel).
Reproduction-ready evidence · last observed 2026-06-06. The source's designated baseline implementation. Reported by source; not independently reproduced.
Current records
Not measured on H100 for this workload. Challenges →
Estimated floor 1.25 µs · record 15703.25× above itestimate, not evidence ›
119.7ms1.00×Reproduction-ready · Apache-2.0 · source2026-06-06
Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.
219.7ms1.00×Reported · Apache-2.0 · source2026-06-06
Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.