submission 69462
Max Brashear · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 84 lines, June 9 Researcher Reciprocity License v1.0.
submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-nvfp4-gemv-69462?include=source"interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp8_e4m3, nvfp4
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:870e19fb056a416d2962d9830ca529cf6e202cb4808a70220de76dd336ed79d1
license declaredunknown
license concludedunknown
authorsMax Brashear
imported2026-08-26
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
fp4
Batched NVFP4(e2m1) GEMV with block FP8 scales.Kernel source
submission.py84 lines
import torch
from task import input_t, output_t
# NVTX for nicer traces during profiling (optional if unavailable)
try:
from torch.cuda.nvtx import range as nvtx_range
except Exception: # pragma: no cover
class nvtx_range: # fallback no-op
def __init__(self, *_args, **_kwargs): pass
def __enter__(self): return self
def __exit__(self, *exc): return False
# ---- helpers (mirrors the reference layout conversion) ---------------------
_SF_VEC = 16 # block size in K used by scaling factors
def _ceil_div(a: int, b: int) -> int:
return (a + b - 1) // b
def _to_blocked(input_matrix: torch.Tensor) -> torch.Tensor:
"""
Reorders a (rows, cols) FP8 scale matrix to the 32x4x...x4x... blocked
format expected by torch._scaled_mm on Hopper/Blackwell parts.
This is kept byte-for-byte compatible with the reference implementation.
"""
rows, cols = input_matrix.shape
n_row_blocks = _ceil_div(rows, 128)
n_col_blocks = _ceil_div(cols, 4)
padded = input_matrix
blocks = padded.view(n_row_blocks, 128, n_col_blocks, 4).permute(0, 2, 1, 3)
rearranged = blocks.reshape(-1, 4, 32, 4).transpose(1, 2).reshape(-1, 32, 16)
return rearranged.flatten()
# ---- main kernel -----------------------------------------------------------
@torch.no_grad()
def custom_kernel(data: input_t) -> output_t:
"""
Batched NVFP4(e2m1) GEMV with block FP8 scales.
Expects a tuple:
(a, b, sfa_ref_cpu, sfb_ref_cpu, _sfa_perm, _sfb_perm, c)
where a:[M,K,L] and b:[1,K,L] are torch.float4_e2m1fn_x2 in K-major order,
sfa/sfb are FP8(e4m3fn) scale tensors (reference layout on CPU),
and c is [M,1,L] in FP16.
For correctness and portability, this implementation mirrors the reference:
for each batch slice l, convert the FP8 scales to the blocked format and
call torch._scaled_mm, then write the single-column result into c.
"""
a_ref, b_ref, sfa_ref_cpu, sfb_ref_cpu, _sfa_perm, _sfb_perm, c_ref = data
M, _, L = a_ref.shape
# Allocate the output (we overwrite every element)
out = torch.empty_like(c_ref)
with nvtx_range("batched_scaled_gemv"):
for l_idx in range(L):
with nvtx_range(f"slice_{l_idx}"):
# Convert the per-slice scales to blocked format (GPU tensors)
scale_a = _to_blocked(sfa_ref_cpu[:, :, l_idx]).cuda()
scale_b = _to_blocked(sfb_ref_cpu[:, :, l_idx]).cuda()
# (M,K) @ (1,K)^T => (M,1); use out_dtype fp16 as required
# b is provided as [1,K] so transpose along K into [K,1]
res = torch._scaled_mm(
a_ref[:, :, l_idx],
b_ref[:, :, l_idx].transpose(0, 1),
scale_a,
scale_b,
bias=None,
out_dtype=torch.float16,
)
# Write result vector into output tensor
out[:, 0, l_idx] = res[:, 0]
return out
scrolls · 84 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON