submission 102763
forestier_kasapi · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 72 lines, June 9 Researcher Reciprocity License v1.0.
submission4.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-nvfp4-gemv-102763?include=source"interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp8_e4m3, nvfp4
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:ba4233194de0068f7266e7bdba9a98025f4feb1a1e0443009a8a79de4280a5f4
license declaredunknown
license concludedunknown
authorsforestier_kasapi
imported2026-08-26
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
fp4
Optimized NVFP4 GEMV using batched preprocessing and vectorized operations.Kernel source
submission4.py72 lines
import torch
from task import input_t, output_t
SF_VEC_SIZE = 16
def to_blocked_batched(input_matrix: torch.Tensor) -> torch.Tensor:
"""
Convert scale-factor matrix to blocked layout for all batches at once.
input_matrix: shape [rows, cols, batch]
Returns: tensor of shape [batch, flattened_size]
"""
rows, cols, batch = input_matrix.shape
n_row_blocks = (rows + 127) >> 7 # Bit shift instead of ceil_div
n_col_blocks = (cols + 3) >> 2
# Single reshape pipeline - minimize intermediate tensors
# [rows, cols, batch] -> [batch, n_row_blocks, 128, n_col_blocks, 4]
blocks = input_matrix.permute(2, 0, 1).reshape(
batch, n_row_blocks, 128, n_col_blocks, 4
)
# [batch, n_row_blocks, n_col_blocks, 128, 4] -> [batch, n_row_blocks*n_col_blocks, 32, 16]
rearranged = (blocks.permute(0, 1, 3, 2, 4)
.reshape(batch, -1, 4, 32, 4)
.permute(0, 1, 3, 2, 4)
.reshape(batch, -1, 32, 16))
# Return flattened batches: [batch, -1]
return rearranged.flatten(1)
def custom_kernel(data: input_t) -> output_t:
"""
Optimized NVFP4 GEMV using batched preprocessing and vectorized operations.
"""
a_ref, b_ref, sfa_ref, sfb_ref, _, _, c_ref = data
_, _, l = c_ref.shape
# Preprocess all scale factors at once (batched conversion)
scales_a = to_blocked_batched(sfa_ref) # [l, -1]
scales_b = to_blocked_batched(sfb_ref) # [l, -1]
# Vectorized batch processing
if l > 1:
# Process multiple batches in parallel when possible
for l_idx in range(l):
a_slice = a_ref[:, :, l_idx]
b_slice = b_ref[:, :, l_idx].transpose(0, 1)
res = torch._scaled_mm(
a_slice,
b_slice,
scales_a[l_idx],
scales_b[l_idx],
bias=None,
out_dtype=torch.float16,
)
c_ref[:, 0, l_idx] = res[:, 0]
else:
# Single batch optimization
res = torch._scaled_mm(
a_ref[:, :, 0],
b_ref[:, :, 0].transpose(0, 1),
scales_a[0],
scales_b[0],
bias=None,
out_dtype=torch.float16,
)
c_ref[:, 0, 0] = res[:, 0]
return c_ref
scrolls · 72 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON