submission 110350
georges314 · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 54 lines, June 9 Researcher Reciprocity License v1.0.
23_perm_scaledmm.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-nvfp4-gemv-110350?include=source"interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp8_e4m3, nvfp4
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:9913039863c4dbfc70123ced9bcced7279a9bff284cddab3cf69452c48685958
license declaredunknown
license concludedunknown
authorsgeorges314
imported2026-08-26
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
fp4
- A, B from data (nvfp4 packed),Kernel source
23_perm_scaledmm.py54 lines
import torch
from task import input_t, output_t
import reference
def custom_kernel(data: input_t) -> output_t:
"""
nvfp4_gemv kernel using:
- A, B from data (nvfp4 packed),
- sfa_permuted / sfb_permuted only for scale reconstruction,
- torch._scaled_mm for the actual GEMV.
This is the cleaned-up version of custom_kernel_from_perm from the notebook.
"""
a_ref, b_ref, sfa_ref, sfb_ref, sfa_perm, sfb_perm, c_ref = data
M, K_half, L = a_ref.shape # logical K = 2*K_half
for l_idx in range(L):
# Permuted scale tensors for this batch
sfa_perm5d = sfa_perm[..., l_idx] # [32, 4, rest_m, 4, rest_k]
sfb_perm5d = sfb_perm[..., l_idx] # [32, 4, rest_n, 4, rest_k]
# Reconstruct flat scales in the exact order torch._scaled_mm expects
scale_a = (
sfa_perm5d
.permute(2, 4, 0, 1, 3) # [rest_m, rest_k, 32, 4, 4]
.contiguous()
.view(-1)
)
scale_b = (
sfb_perm5d
.permute(2, 4, 0, 1, 3)
.contiguous()
.view(-1)
)
# A: [m, k_half] (nvfp4 packed)
# B: [1..128, k_half] (nvfp4 packed) -> transpose for _scaled_mm: [k_half, n]
A_mat = a_ref[:, :, l_idx] # [M, K_half]
B_mat = b_ref[:, :, l_idx].transpose(0, 1) # [K_half, n]
res = torch._scaled_mm(
A_mat,
B_mat,
scale_a,
scale_b,
bias=None,
out_dtype=torch.float16,
)
c_ref[:, 0, l_idx] = res[:, 0]
return c_ref
scrolls · 54 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON