Skip to content
KernelIndex
Search⌘K

submission 103427

Founzo · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 54 lines, June 9 Researcher Reciprocity License v1.0.

submission_v3.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-nvfp4-gemv-103427?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp8_e4m3, nvfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
NVFP4 GEMVsuite of 3 cases
NVIDIA B200
64.4µs
#305 of 678
2025-11-25

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:29965f12b97fa0bd9b41990832eb34d821acb2b549dfca17ed91c97a609b9bbc
license declaredunknown
license concludedunknown
authorsFounzo
imported2026-08-26

Kernel source

submission_v3.py54 lines
import torch
from task import input_t, output_t

# Kernel configuration parameters
sf_vec_size = 16

def scale_from_permuted(sf_perm, l_idx):
    sf = sf_perm[..., l_idx]
    sf = sf.permute(2, 4, 0, 1, 3)
    sf = sf.reshape(-1, 32, 16)
    return sf.flatten()

def custom_kernel(
    data: input_t,
) -> output_t:
    """
    m: Number of rows in matrix A
    k: Number of columns in A (and length of vector b)
    l: Batch size
    Args:
        data: Tuple that expands to:
            a: [m, k, l] - Input matrix in torch.float4e2m1fn_x2 data type,
            b: [1, k, l] - Input vector in torch.float4e2m1fn_x2 data type,
            scale_a: [m, k, l] - Input scale factors in torch.float8e4m3fn data type,
            scale_b: [1, k, l] - Input scale factors in torch.float8e4m3fn data type,
            scale_a_permuted: [32, 4, rest_m, 4, rest_k, l] - Input scale factors in torch.float8e4m3fn data type,
            scale_b_permuted: [32, 4, rest_n, 4, rest_k, l] - Input scale factors in torch.float8e4m3fn data type,
            c: [m, 1, l] - Output vector in torch.float16 data type
    Returns:
        c: [m, 1, l] - Output vector in torch.float16 data type
    """
    
    a, b, sfa_cpu, sfb_cpu, sfa_perm, sfb_perm, c = data

    _, _, l = c.shape

    for l_idx in range(l):
        scale_a = scale_from_permuted(sfa_perm, l_idx)
        scale_b = scale_from_permuted(sfb_perm, l_idx)

        # GEMV bloqué : (m, k) @ (k, 1) -> (m, 1)
        res = torch._scaled_mm(
            a[:, :, l_idx], # (m, k)
            b[:, :, l_idx].transpose(0, 1), # (k, 1)
            scale_a, 
            scale_b,
            bias=None,
            out_dtype=torch.float16,
        )

        c[:, 0, l_idx] = res[:, 0]

    return c
scrolls · 54 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON