Skip to content
KernelIndex
Search⌘K

submission 456302

Nick Nuon · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 90 lines, June 9 Researcher Reciprocity License v1.0.

control.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-nvfp4-group-gemm-456302?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp8_e4m3, nvfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
NVFP4 group GEMMsuite of 4 cases
NVIDIA B200
8.68ms
#303 of 310
2026-02-04

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:512c74ff13bdef093ae1ecbd25f47bd726f57b04d3849bce0951caf06d74de32
license declaredunknown
license concludedunknown
authorsNick Nuon
imported2026-08-15

Kernel source

control.py90 lines
#!/usr/bin/env python3
import torch

# ## Benchmarks:
# ```
# g: 8; k: [7168, 7168, 7168, 7168, 7168, 7168, 7168, 7168]; m: [80, 176, 128, 72, 64, 248, 96, 160]; n: [4096, 4096, 4096, 4096, 4096, 4096, 4096, 4096]; seed: 1111
#  ⏱ 37.2 ± 0.04 ms
#  ⚡ 36.9 ms 🐌 37.5 ms

# g: 8; k: [2048, 2048, 2048, 2048, 2048, 2048, 2048, 2048]; m: [40, 76, 168, 72, 164, 148, 196, 160]; n: [7168, 7168, 7168, 7168, 7168, 7168, 7168, 7168]; seed: 1111
#  ⏱ 18.7 ± 0.03 ms
#  ⚡ 18.3 ms 🐌 20.3 ms

# g: 2; k: [4096, 4096]; m: [192, 320]; n: [3072, 3072]; seed: 1111
#  ⏱ 4.15 ± 0.004 ms
#  ⚡ 4.09 ms 🐌 4.22 ms

# g: 2; k: [1536, 1536]; m: [128, 384]; n: [4096, 4096]; seed: 1111
#  ⏱ 2.03 ± 0.004 ms
#  ⚡ 1964 µs 🐌 2.14 ms

#     Personal best on NVIDIA: 45.7 µs

# ----------------------------
# Helpers (same as reference)
# ----------------------------
sf_vec_size = 16

def ceil_div(a, b):
    return (a + b - 1) // b

def to_blocked(input_matrix):
    rows, cols = input_matrix.shape

    n_row_blocks = ceil_div(rows, 128)
    n_col_blocks = ceil_div(cols, 4)
    padded_rows = n_row_blocks * 128
    padded_cols = n_col_blocks * 4

    if padded_rows != rows or padded_cols != cols:
        padded = torch.nn.functional.pad(
            input_matrix,
            (0, padded_cols - cols, 0, padded_rows - rows),
            mode="constant",
            value=0,
        )
    else:
        padded = input_matrix

    blocks = padded.view(n_row_blocks, 128, n_col_blocks, 4).permute(0, 2, 1, 3)
    rearranged = blocks.reshape(-1, 4, 32, 4).transpose(1, 2).reshape(-1, 32, 16)
    return rearranged.flatten()

# ----------------------------
# GPU Mode entrypoint
# ----------------------------
@torch.inference_mode()
def custom_kernel(data):
    abc_tensors, sfasfb_tensors, _reordered, problem_sizes = data

    outs = []
    for (a, b, c), (sfa, sfb), (m, n, k, l) in zip(
        abc_tensors, sfasfb_tensors, problem_sizes
    ):
        # c is already cuda in the generator, but keep it safe
        if c.device.type != "cuda":
            c = c.cuda()

        for l_idx in range(l):
            # scale factors -> blocked layout on GPU
            scale_a = to_blocked(sfa[:, :, l_idx]).to("cuda", non_blocking=True)
            scale_b = to_blocked(sfb[:, :, l_idx]).to("cuda", non_blocking=True)

            # IMPORTANT:
            # - A slice is already float4_e2m1fn_x2
            # - Bt must be a transpose VIEW, NOT contiguous()
            A  = a[:, :, l_idx]                  # dtype float4_e2m1fn_x2
            Bt = b[:, :, l_idx].transpose(0, 1)  # no contiguous(), no copy

            c[:, :, l_idx] = torch._scaled_mm(
                A, Bt,
                scale_a, scale_b,
                bias=None,
                out_dtype=torch.float16,
            )

        outs.append(c)

    return outs
scrolls · 90 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON