Skip to content
KernelIndex
Search⌘K

submission 69323

gau.nernst · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 34 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-nvfp4-gemv-69323?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp8_e4m3, nvfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
NVFP4 GEMVsuite of 3 cases
NVIDIA B200
55.4µs
#286 of 678
2025-11-10

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:f984907d40e1ffd2d56931e3abddc3e70b2f92b78f8eb647a17c4ddecacc3cdf
license declaredunknown
license concludedunknown
authorsgau.nernst
imported2026-08-15

Kernel source

submission.py34 lines
#!POPCORN leaderboard nvfp4_gemv

import torch
from task import input_t, output_t


def custom_kernel(data: input_t) -> output_t:
    # a:   [  M, K, L],                   natural shape [L,   M, K]
    # b:   [128, K, L],                   natural shape [L, 128, K] - only the 1st row is used
    # sfa: [32, 4, rest_m, 4, rest_k, L], natural shape [L, rest_m, rest_k, 32, 4, 4]
    # sfb: [32, 4,      1, 4, rest_k, L], natural shape [L,      1, rest_k, 32, 4, 4]
    # c:   [  M, 1, L],                   natural shape [L, M, 1]
    a, b, _, _, sfa, sfb, c_ref = data
    M, _, L = c_ref.shape

    a = a.permute(2, 0, 1)  # [L, M, K/2]
    b = b.permute(2, 0, 1)  # [L, 128, K/2]
    sfa = sfa.permute(5, 2, 4, 0, 1, 3).view(L, M, -1)
    sfb = sfb.permute(5, 2, 4, 0, 1, 3).view(L, 128, -1)

    big_c = c_ref.new_empty(L, 128, M).transpose(1, 2)  # (L, M, 128), M-major

    for l_idx in range(L):
        torch._scaled_mm(
            a[l_idx],
            b[l_idx].transpose(0, 1),
            sfa[l_idx],
            sfb[l_idx],
            out_dtype=torch.float16,
            out=big_c[l_idx],
        )

    return big_c[..., :1].permute(1, 2, 0)  # convert to [L, M, 1]
scrolls · 34 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON