submission 117057
Founzo · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 53 lines, June 9 Researcher Reciprocity License v1.0.
submission_v1.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-nvfp4-gemm-117057?include=source"interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp8_e4m3, nvfp4
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:63df9e0450606ae38e27c66fa2b0ddd3d3da97312255c585c0086bcced454d30
license declaredunknown
license concludedunknown
authorsFounzo
imported2026-08-26
Kernel source
submission_v1.py53 lines
import torch
from task import input_t, output_t
# Kernel configuration parameters
sf_vec_size = 16
def scale_from_permuted(sf_perm, l_idx):
sf = sf_perm[..., l_idx]
sf = sf.permute(2, 4, 0, 1, 3)
sf = sf.reshape(-1, 32, 16)
return sf.flatten()
def custom_kernel(
data: input_t,
) -> output_t:
"""
m: Number of rows in matrix A
k: Number of columns in A (and length of vector b)
l: Batch size
Args:
data: Tuple that expands to:
a: [m, k, l] - Input matrix in torch.float4e2m1fn_x2 data type,
b: [1, k, l] - Input vector in torch.float4e2m1fn_x2 data type,
scale_a: [m, k, l] - Input scale factors in torch.float8e4m3fn data type,
scale_b: [1, k, l] - Input scale factors in torch.float8e4m3fn data type,
scale_a_permuted: [32, 4, rest_m, 4, rest_k, l] - Input scale factors in torch.float8e4m3fn data type,
scale_b_permuted: [32, 4, rest_n, 4, rest_k, l] - Input scale factors in torch.float8e4m3fn data type,
c: [m, n, l] - Output matrix in torch.float16 data type
Returns:
c: [m, n, l] - Output matrix in torch.float16 data type
"""
a, b, sfa_cpu, sfb_cpu, sfa_perm, sfb_perm, c = data
_, _, l = c.shape
scale_a = scale_from_permuted(sfa_perm, 0)
scale_b = scale_from_permuted(sfb_perm, 0)
# GEMV bloqué : (m, k) @ (k, n) -> (m, n)
res = torch._scaled_mm(
a[:, :, 0], # (m, k)
b[:, :, 0].transpose(0, 1), # (k, n)
scale_a, # (-1, 32, 16)
scale_b, # (-1, 32, 16)
bias=None,
out_dtype=torch.float16,
)
c[:, :, 0] = res
return c
scrolls · 53 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON