Grouped GEMM FP8 fp4 m masked g4 n4096 k2048
M-masked grouped mixed FP8xFP4 GEMM (4 groups, max_m=512, N=4096, K=2048), DeepGEMM SM100. Each group g computes A_g[:masked_m[g]] @ B_g.T into a fixed [max_m] bucket; rows >= masked_m[g] are zeroed. A is FP8 (float8_e4m3fn, per-token x128 float32 scale); B per-group FP4 (E2M1 int8 + UE8M0 per-32 float32 scale). expected_m is the scheduling hint (int(1.2*expected_per_group)). Reference dequantizes per group; kernel is exact.
Reported evidence · last observed 2026-06-06. The source's designated baseline implementation. Reported by source; not independently reproduced.
Current records
Not measured on H100 for this workload. Challenges →
Estimated floor 3.16 µs · record 50.71× above itestimate, not evidence ›
1160.3µs1.00×Reported · Apache-2.0 · source2026-06-06
Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.
2162.1µs1.01×Reported · Apache-2.0 · source2026-06-06
Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.