Skip to content
KernelIndex
Search⌘K

Grouped GEMM NVFP4 m contiguous g4 n2048 k2048

2 eligible runs
gemm

M-contiguous grouped NVFP4 GEMM (4 groups, N=2048, K=2048). Rows of all groups concatenated; m_indptr gives group boundaries. NVFP4: packed E2M1 (int8) + UE4M3 per-16 block scale (logical layout, int8) + per-group global de-scale alpha. Hand-written CUDA baseline (runs on sm100, unlike the sm120-only flashinfer grouped kernel).

Source baseline · unbeaten
19.7msmean
cuda_nvfp4_grouped_naive_g4_n2048_k2048FlashInfer-Bench baselines · Apache-2.0 · cuda

Reproduction-ready evidence · last observed 2026-06-06. The source's designated baseline implementation. Reported by source; not independently reproduced.

Current records

Not measured on H100 for this workload. Challenges →

Source-native comparison · GPU NVIDIA B200 · Workload m = 512 · fp32 · CUDA 13.0 · Framework pytorch 2.11.0+cu130 · Protocol flashinfer-bench · mean · 2 results · last observed 2026-06-06Record history →
Estimated floor 1.25 µs · record 15703.25× above itestimate, not evidence ›
DRAM 1.25 µs · bandwidth-bound on B200
every declared tensor crosses HBM exactly once (8,000 GB/s, B200 datasheet)
no arithmetic formula for this family: bandwidth floor only
headroom-v1: a lower bound from declared tensors and datasheet peaks. A kernel can sit well above it for good reasons.
#
Implementation
Latency
vs #1
Trust
Observed
1
cuda_nvfp4_grouped_naive_g4_n2048_k2048baselineFlashInfer-Bench baselines
19.7ms
1.00×
Reproduction-ready · Apache-2.0 · source
2026-06-06

Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.

source mirroredApache-2.0no install recipeView source →Run detail →
2
cuda_nvfp4_grouped_naive_g4_n2048_k2048baselineFlashInfer-Bench baselines
19.7ms
1.00×
Reported · Apache-2.0 · source
2026-06-06

Measured exactly what you asked. The source's designated baseline implementation. Reported by source; not independently reproduced.

source mirroredApache-2.0no install recipeView source →Run detail →

Implementations

Implementation
Runtime
Best latency
Evidence
Availability
cuda_nvfp4_grouped_naive_g4_n2048_k2048FlashInfer-Bench baselines
cuda
19.7ms
1.00×
Reproduction-ready
Apache-2.0 · source

Semantics

Inputs and outputs
a_fp4int8 [m, k_half]
a_scaleint8 [m, k_blocks]
b_fp4int8 [g, n, k_half]
b_scaleint8 [g, n, k_blocks]
m_indptrint32 [gp1]
alphafp32 [g]
cbf16 [m, n]
Axes and behavior
gconstant = 4
kconstant = 2048
mvariable
nconstant = 2048
gp1constant = 5
k_halfconstant = 1024
k_blocksconstant = 128
determinismunspecified
constraintsNo mutation or aliasing
Identity
aliasgrouped_gemm_nvfp4_m_contiguous_g4_n2048_k2048sha2564b7c8c34f989…
Sources: FlashInfer-Bench (2026-06-06) · Apache-2.0last observed 2026-06-06How records are decidedJSON