Skip to content
KernelIndex
Search⌘K

submission 345025

zyn · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 90 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-nvfp4-dual-gemm-345025?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp8_e4m3, nvfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
NVFP4 dual GEMMsuite of 4 cases
NVIDIA B200
89.7µs
#407 of 420
2026-01-14

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:5c92ec83a63918ce7d0503a9d87bab8cf8640ccafa16109f292c0d2e70ab9523
license declaredunknown
license concludedunknown
authorszyn
imported2026-08-26

Kernel source

submission.py90 lines
import torch
from task import input_t, output_t
from utils import make_match_reference

sf_vec_size = 16


def ceil_div(a, b):
    return (a + b - 1) // b


def to_blocked(input_matrix):
    rows, cols = input_matrix.shape
    n_row_blocks = ceil_div(rows, 128)
    n_col_blocks = ceil_div(cols, 4)
    blocks = input_matrix.view(
        n_row_blocks, 128, n_col_blocks, 4
    ).permute(0, 2, 1, 3)
    rearranged = (
        blocks.reshape(-1, 4, 32, 4)
        .transpose(1, 2)
        .reshape(-1, 32, 16)
    )
    return rearranged.flatten()


def custom_kernel(data: input_t) -> output_t:
    """
    Slightly optimized reference implementation:
    - precompute blocked scales
    - reuse transposed B
    - minimize Python/GPU overhead
    """
    (
        a,
        b1,
        b2,
        sfa_cpu,
        sfb1_cpu,
        sfb2_cpu,
        _,
        _,
        _,
        c_ref,
    ) = data

    m, n, l = c_ref.shape

    # ---- precompute blocked scale factors (once) ----
    scale_a = [
        to_blocked(sfa_cpu[:, :, li]).cuda() for li in range(l)
    ]
    scale_b1 = [
        to_blocked(sfb1_cpu[:, :, li]).cuda() for li in range(l)
    ]
    scale_b2 = [
        to_blocked(sfb2_cpu[:, :, li]).cuda() for li in range(l)
    ]

    # ---- precompute transposed B ----
    b1_t = b1.transpose(0, 1)
    b2_t = b2.transpose(0, 1)

    out1 = torch.empty_like(c_ref, dtype=torch.float32)
    out2 = torch.empty_like(c_ref, dtype=torch.float32)

    for li in range(l):
        out1[:, :, li] = torch._scaled_mm(
            a[:, :, li],
            b1_t[:, :, li],
            scale_a[li],
            scale_b1[li],
            out_dtype=torch.float32,
        )

        out2[:, :, li] = torch._scaled_mm(
            a[:, :, li],
            b2_t[:, :, li],
            scale_a[li],
            scale_b2[li],
            out_dtype=torch.float32,
        )

    return (torch.nn.functional.silu(out1) * out2).to(torch.float16)


check_implementation = make_match_reference(
    custom_kernel, rtol=1e-3, atol=1e-3
)
scrolls · 90 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON