submission 456302
Nick Nuon · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 90 lines, June 9 Researcher Reciprocity License v1.0.
control.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-nvfp4-group-gemm-456302?include=source"interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp8_e4m3, nvfp4
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:512c74ff13bdef093ae1ecbd25f47bd726f57b04d3849bce0951caf06d74de32
license declaredunknown
license concludedunknown
authorsNick Nuon
imported2026-08-15
Kernel source
control.py90 lines
#!/usr/bin/env python3
import torch
# ## Benchmarks:
# ```
# g: 8; k: [7168, 7168, 7168, 7168, 7168, 7168, 7168, 7168]; m: [80, 176, 128, 72, 64, 248, 96, 160]; n: [4096, 4096, 4096, 4096, 4096, 4096, 4096, 4096]; seed: 1111
# ⏱ 37.2 ± 0.04 ms
# ⚡ 36.9 ms 🐌 37.5 ms
# g: 8; k: [2048, 2048, 2048, 2048, 2048, 2048, 2048, 2048]; m: [40, 76, 168, 72, 164, 148, 196, 160]; n: [7168, 7168, 7168, 7168, 7168, 7168, 7168, 7168]; seed: 1111
# ⏱ 18.7 ± 0.03 ms
# ⚡ 18.3 ms 🐌 20.3 ms
# g: 2; k: [4096, 4096]; m: [192, 320]; n: [3072, 3072]; seed: 1111
# ⏱ 4.15 ± 0.004 ms
# ⚡ 4.09 ms 🐌 4.22 ms
# g: 2; k: [1536, 1536]; m: [128, 384]; n: [4096, 4096]; seed: 1111
# ⏱ 2.03 ± 0.004 ms
# ⚡ 1964 µs 🐌 2.14 ms
# Personal best on NVIDIA: 45.7 µs
# ----------------------------
# Helpers (same as reference)
# ----------------------------
sf_vec_size = 16
def ceil_div(a, b):
return (a + b - 1) // b
def to_blocked(input_matrix):
rows, cols = input_matrix.shape
n_row_blocks = ceil_div(rows, 128)
n_col_blocks = ceil_div(cols, 4)
padded_rows = n_row_blocks * 128
padded_cols = n_col_blocks * 4
if padded_rows != rows or padded_cols != cols:
padded = torch.nn.functional.pad(
input_matrix,
(0, padded_cols - cols, 0, padded_rows - rows),
mode="constant",
value=0,
)
else:
padded = input_matrix
blocks = padded.view(n_row_blocks, 128, n_col_blocks, 4).permute(0, 2, 1, 3)
rearranged = blocks.reshape(-1, 4, 32, 4).transpose(1, 2).reshape(-1, 32, 16)
return rearranged.flatten()
# ----------------------------
# GPU Mode entrypoint
# ----------------------------
@torch.inference_mode()
def custom_kernel(data):
abc_tensors, sfasfb_tensors, _reordered, problem_sizes = data
outs = []
for (a, b, c), (sfa, sfb), (m, n, k, l) in zip(
abc_tensors, sfasfb_tensors, problem_sizes
):
# c is already cuda in the generator, but keep it safe
if c.device.type != "cuda":
c = c.cuda()
for l_idx in range(l):
# scale factors -> blocked layout on GPU
scale_a = to_blocked(sfa[:, :, l_idx]).to("cuda", non_blocking=True)
scale_b = to_blocked(sfb[:, :, l_idx]).to("cuda", non_blocking=True)
# IMPORTANT:
# - A slice is already float4_e2m1fn_x2
# - Bt must be a transpose VIEW, NOT contiguous()
A = a[:, :, l_idx] # dtype float4_e2m1fn_x2
Bt = b[:, :, l_idx].transpose(0, 1) # no contiguous(), no copy
c[:, :, l_idx] = torch._scaled_mm(
A, Bt,
scale_a, scale_b,
bias=None,
out_dtype=torch.float16,
)
outs.append(c)
return outs
scrolls · 90 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON