submission 75968
arsrivish26691 · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 107 lines, June 9 Researcher Reciprocity License v1.0.
nvfp4tsk1vg.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-nvfp4-gemv-75968?include=source"interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp8_e4m3, nvfp4
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:16a582313332c2a5b26c98f7c1de98a826ed5d044df69d0cdcaa9c9a10617653
license declaredunknown
license concludedunknown
authorsarsrivish26691
imported2026-08-26
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
fp4
NVFP4 block-scaled GEMV using torch._scaled_mm, with:Kernel source
nvfp4tsk1vg.py107 lines
import torch
from task import input_t, output_t
def ceil_div(a: int, b: int) -> int:
return (a + b - 1) // b
def to_blocked_all_gpu(sf: torch.Tensor) -> torch.Tensor:
"""
Batched, GPU version of the provided `to_blocked`.
Input:
sf: [M, K_sf, L] (fp8) on CUDA, with:
- M multiple of 128
- K_sf multiple of 4
Output:
scales_all: [L, num_scales_per_matrix]
where scales_all[l] == to_blocked(sf[:, :, l]) from the reference.
"""
assert sf.dim() == 3, f"Expected 3D scale tensor, got {sf.shape}"
M, K_sf, L = sf.shape
assert M % 128 == 0, f"M={M} must be divisible by 128"
assert K_sf % 4 == 0, f"K_sf={K_sf} must be divisible by 4"
n_row_blocks = M // 128
n_col_blocks = K_sf // 4
# Start from [M, K_sf, L] on GPU
# 1) [M, K_sf, L] -> [n_row_blocks, 128, n_col_blocks, 4, L]
blocks = sf.view(n_row_blocks, 128, n_col_blocks, 4, L)
# 2) [n_row_blocks, n_col_blocks, 128, 4, L]
blocks = blocks.permute(0, 2, 1, 3, 4)
# 3) merge row-block dims: [B, 4, 32, 4, L], B = n_row_blocks * n_col_blocks
blocks = blocks.reshape(-1, 4, 32, 4, L)
# 4) transpose 4 x 32 -> 32 x 4: [B, 32, 4, 4, L]
blocks = blocks.transpose(1, 2)
# 5) merge the two 4s -> 16: [B, 32, 16, L]
blocks = blocks.reshape(-1, 32, 16, L)
# 6) move L in front, flatten per-matrix: [L, B, 32, 16] -> [L, B*32*16]
blocks = blocks.permute(3, 0, 1, 2).contiguous() # [L, B, 32, 16]
L_, B, _, _ = blocks.shape
return blocks.view(L_, B * 32 * 16) # [L, num_scales]
def custom_kernel(data: input_t) -> output_t:
"""
NVFP4 block-scaled GEMV using torch._scaled_mm, with:
- scale factors moved to GPU once
- scale blocking/swizzling done on GPU for all L at once
- inner loop over L only runs `_scaled_mm` and a copy
"""
a_ref, b_ref, sfa_ref_cpu, sfb_ref_cpu, _, _, c_ref = data
device = a_ref.device # should be "cuda"
M, _, L = c_ref.shape
# ---- Move scale tensors to GPU once ----
# sfa_ref_cpu : [M, K//16, L] on CPU → GPU
# sfb_ref_cpu : [128, K//16, L] on CPU → GPU
sfa_gpu = sfa_ref_cpu.to(device=device, non_blocking=True)
sfb_gpu = sfb_ref_cpu.to(device=device, non_blocking=True)
# ---- Precompute blocked scale vectors for ALL L on GPU ----
# These are equivalent to:
# for l in range(L): to_blocked(sfa_ref_cpu[:, :, l])
scale_a_all = to_blocked_all_gpu(sfa_gpu) # [L, num_scales_a] on GPU
scale_b_all = to_blocked_all_gpu(sfb_gpu) # [L, num_scales_b] on GPU
# ---- Main GEMV loop over batch L ----
# We still need one `_scaled_mm` per l (batches are independent),
# but now there is:
# - no CPU swizzle work
# - no per-l .cuda() calls for scale vectors
for l_idx in range(L):
# A_l: [M, K]
a_l = a_ref[:, :, l_idx]
# B_l: original is [128, K/2, L] in nvfp4; reference uses transpose
# b_ref[:, :, l].T: [K, 128]
b_l = b_ref[:, :, l_idx].transpose(0, 1)
sa = scale_a_all[l_idx] # [num_scales_a] on GPU
sb = scale_b_all[l_idx] # [num_scales_b] on GPU
# Compute (M, K) @ (K, 128) with NVFP4 + block scales
res = torch._scaled_mm(
a_l,
b_l,
sa,
sb,
bias=None,
out_dtype=torch.float16,
) # [M, 128]
# c_ref: [M, 1, L] – GEMV, real N=1, we only keep the first column
c_ref[:, 0, l_idx].copy_(res[:, 0])
return c_ref
scrolls · 107 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON