Skip to content
KernelIndex
Search⌘K

submission 114372

peisong · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 50 lines, June 9 Researcher Reciprocity License v1.0.

sol-1.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-nvfp4-gemm-114372?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp8_e4m3, nvfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
NVFP4 GEMMsuite of 3 cases
NVIDIA B200
50.5µs
#282 of 369
2025-11-29

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:100cdd8cffcd5cc0f43a1caa06e849585e9aa97d1faf61d97db5f401669cb23e
license declaredunknown
license concludedunknown
authorspeisong
imported2026-08-26

Kernel source

sol-1.py50 lines
"""Optimized reference kernel using torch._scaled_mm on GPU-resident data.

This mirrors the interface in ``template.py`` but avoids any CPU-side scale
factor reformatting by performing the block layout conversion directly on the
GPU before calling ``torch._scaled_mm``.
"""

import torch
from task import input_t, output_t


sf_vec_size = 16


def ceil_div(a: int, b: int) -> int:
    return (a + b - 1) // b


def _to_blocked_gpu(scale: torch.Tensor) -> torch.Tensor:
    """Convert (rows, cols) scale tensor to blocked layout expected by scaled_mm."""
    rows, cols = scale.shape
    n_row_blocks = ceil_div(rows, 128)
    n_col_blocks = ceil_div(cols, 4)
    blocks = scale.view(n_row_blocks, 128, n_col_blocks, 4).permute(0, 2, 1, 3)
    rearranged = blocks.reshape(-1, 4, 32, 4).transpose(1, 2).reshape(-1, 32, 16)
    return rearranged.flatten()


def custom_kernel(data: input_t) -> output_t:
    """Reference-style implementation using Torch scaled GEMM on GPU."""
    a, b, sfa, sfb, _, _, c = data

    _, _, l = c.shape
    for l_idx in range(l):
        scale_a = _to_blocked_gpu(sfa[:, :, l_idx])
        scale_b = _to_blocked_gpu(sfb[:, :, l_idx])
        c[:, :, l_idx] = torch._scaled_mm(
            a[:, :, l_idx],
            b[:, :, l_idx].transpose(0, 1),
            scale_a,
            scale_b,
            bias=None,
            out_dtype=torch.float16,
        )

    return c


__all__ = ["custom_kernel"]
scrolls · 50 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON