Skip to content
KernelIndex
Search⌘K

submission 487985

nataliakokoromyti · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 45 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-nvfp4-group-gemm-487985?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp8_e4m3, nvfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
NVFP4 group GEMMsuite of 4 cases
NVIDIA B200
75.0µs
#81 of 145
2026-02-10

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:4c43fc44ba0539bda457ff5f96e3bc43674cf42d10b4b6aa794c87f5a5267941
license declaredunknown
license concludedunknown
authorsnataliakokoromyti
imported2026-08-15

Kernel source

submission.py45 lines
import torch
from task import input_t, output_t


@torch.no_grad()
def custom_kernel(data: input_t) -> output_t:
    abc_tensors, sfasfb_tensors, sfasfb_reordered_tensors, problem_sizes = data
    num_groups = len(abc_tensors)
    result_tensors = []

    for i in range(num_groups):
        a_ref, b_ref, c_ref = abc_tensors[i]
        sfa_reordered, sfb_reordered = sfasfb_reordered_tensors[i]
        m, n, k, l = problem_sizes[i]

        # Reordered shape: (32, 4, n_rb, 4, n_cb, L)
        # Target blocked order matching to_blocked() output.
        scale_a = (
            sfa_reordered.permute(2, 4, 0, 1, 3, 5)
            [:, :, :, :, :, 0]
            .contiguous()
            .reshape(-1, 32, 16)
            .flatten()
        )
        scale_b = (
            sfb_reordered.permute(2, 4, 0, 1, 3, 5)
            [:, :, :, :, :, 0]
            .contiguous()
            .reshape(-1, 32, 16)
            .flatten()
        )

        res = torch._scaled_mm(
            a_ref[:, :, 0].view(torch.float4_e2m1fn_x2),
            b_ref[:, :, 0].transpose(0, 1).view(torch.float4_e2m1fn_x2),
            scale_a,
            scale_b,
            bias=None,
            out_dtype=torch.float16,
        )
        c_ref[:, :, 0] = res
        result_tensors.append(c_ref)

    return result_tensors
scrolls · 45 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Changes from previous submission

Against this author's previous submission submission 400223.

⋯ diff truncated: revisions differ almost entirely

Best evidence level for this revision: reported

JSON