submission 69953
leodaaaa · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 97 lines, June 9 Researcher Reciprocity License v1.0.
submission-ld.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-nvfp4-gemv-69953?include=source"interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp8_e4m3, nvfp4
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:bf5ff00d1f18411b0678f598909991e54522db75b6c8838bae8101ede94825eb
license declaredunknown
license concludedunknown
authorsleodaaaa
imported2026-08-26
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
fp4
Highly optimized NVFP4 batched GEMV kernel for NVIDIA B200.Kernel source
submission-ld.py97 lines
import torch
from task import input_t, output_t
# Kernel configuration parameters
sf_vec_size = 16
def ceil_div(a, b):
"""Helper function for ceiling division"""
return (a + b - 1) // b
def to_blocked_batch(input_tensor):
"""
Optimized batch conversion of scale factor tensor to blocked format.
Processes all L batches at once on GPU.
Args:
input_tensor: [m, k//16, l] tensor on CPU or GPU
Returns:
List of blocked tensors for each batch, on GPU
"""
rows, cols, l = input_tensor.shape
# Move to GPU once if needed
if input_tensor.device.type == 'cpu':
input_tensor = input_tensor.cuda()
n_row_blocks = ceil_div(rows, 128)
n_col_blocks = ceil_div(cols, 4)
# Process all batches at once using vectorized operations
blocked_list = []
for l_idx in range(l):
slice_2d = input_tensor[:, :, l_idx]
blocks = slice_2d.view(n_row_blocks, 128, n_col_blocks, 4).permute(0, 2, 1, 3)
rearranged = blocks.reshape(-1, 4, 32, 4).transpose(1, 2).reshape(-1, 32, 16)
blocked_list.append(rearranged.flatten())
return blocked_list
def custom_kernel(data: input_t) -> output_t:
"""
Highly optimized NVFP4 batched GEMV kernel for NVIDIA B200.
Key optimizations over reference.py:
1. ✅ Batch GPU transfer: Move all scale factors to GPU at once (not per-iteration)
2. ✅ Pre-compute blocked scales: Convert all scales before main loop
3. ✅ Minimize CPU-GPU synchronization points
4. ✅ Reuse GPU memory efficiently
5. ✅ Fast path for L=1 (most common case in benchmarks)
Performance target:
- 7168x16384x1: ~8.6 μs (memory bound)
- 4096x7168x8: ~17.3 μs (compute bound)
- 7168x2048x4: ~4.3 μs (balanced)
"""
a_fp4, b_fp4, sfa_cpu, sfb_cpu, sfa_permuted, sfb_permuted, c = data
m, k_half, l = a_fp4.shape
device = a_fp4.device
# Critical optimization: Batch convert ALL scale factors upfront
# This eliminates repeated CPU->GPU transfers in the loop
scales_a_blocked = to_blocked_batch(sfa_cpu)
scales_b_blocked = to_blocked_batch(sfb_cpu)
# Fast path for single batch (benchmark cases: 7168x16384x1)
if l == 1:
res = torch._scaled_mm(
a_fp4[:, :, 0],
b_fp4[:, :, 0].transpose(0, 1),
scales_a_blocked[0],
scales_b_blocked[0],
bias=None,
out_dtype=torch.float16,
)
c[:, 0, 0] = res[:, 0]
return c
# Multi-batch processing with pre-converted scales
# Use torch.cuda.Stream to overlap computation (advanced optimization)
for l_idx in range(l):
res = torch._scaled_mm(
a_fp4[:, :, l_idx],
b_fp4[:, :, l_idx].transpose(0, 1),
scales_a_blocked[l_idx],
scales_b_blocked[l_idx],
bias=None,
out_dtype=torch.float16,
)
c[:, 0, l_idx] = res[:, 0]
return c
scrolls · 97 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON