submission 114372
peisong · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 50 lines, June 9 Researcher Reciprocity License v1.0.
sol-1.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-nvfp4-gemm-114372?include=source"interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp8_e4m3, nvfp4
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:100cdd8cffcd5cc0f43a1caa06e849585e9aa97d1faf61d97db5f401669cb23e
license declaredunknown
license concludedunknown
authorspeisong
imported2026-08-26
Kernel source
sol-1.py50 lines
"""Optimized reference kernel using torch._scaled_mm on GPU-resident data.
This mirrors the interface in ``template.py`` but avoids any CPU-side scale
factor reformatting by performing the block layout conversion directly on the
GPU before calling ``torch._scaled_mm``.
"""
import torch
from task import input_t, output_t
sf_vec_size = 16
def ceil_div(a: int, b: int) -> int:
return (a + b - 1) // b
def _to_blocked_gpu(scale: torch.Tensor) -> torch.Tensor:
"""Convert (rows, cols) scale tensor to blocked layout expected by scaled_mm."""
rows, cols = scale.shape
n_row_blocks = ceil_div(rows, 128)
n_col_blocks = ceil_div(cols, 4)
blocks = scale.view(n_row_blocks, 128, n_col_blocks, 4).permute(0, 2, 1, 3)
rearranged = blocks.reshape(-1, 4, 32, 4).transpose(1, 2).reshape(-1, 32, 16)
return rearranged.flatten()
def custom_kernel(data: input_t) -> output_t:
"""Reference-style implementation using Torch scaled GEMM on GPU."""
a, b, sfa, sfb, _, _, c = data
_, _, l = c.shape
for l_idx in range(l):
scale_a = _to_blocked_gpu(sfa[:, :, l_idx])
scale_b = _to_blocked_gpu(sfb[:, :, l_idx])
c[:, :, l_idx] = torch._scaled_mm(
a[:, :, l_idx],
b[:, :, l_idx].transpose(0, 1),
scale_a,
scale_b,
bias=None,
out_dtype=torch.float16,
)
return c
__all__ = ["custom_kernel"]
scrolls · 50 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON