submission 105294
fl4res · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 78 lines, June 9 Researcher Reciprocity License v1.0.
fp4_nogcc.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-nvfp4-gemv-105294?include=source"interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp8_e4m3, nvfp4
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:491b44f8c2502727086010a794a3c37ab5959b6f884b792619510680f7565f0d
license declaredunknown
license concludedunknown
authorsfl4res
imported2026-08-26
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
fp4
Simple optimized NVFP4 block-scaled GEMV kernel.Kernel source
fp4_nogcc.py78 lines
"""
Simple optimized NVFP4 block-scaled GEMV kernel.
No CUDA graphs - just pure vectorized operations.
"""
import torch
from task import input_t, output_t
def ceil_div(a: int, b: int) -> int:
return (a + b - 1) // b
@torch.no_grad()
def to_blocked_all(sf: torch.Tensor) -> list:
"""
Convert (rows, cols, L) scale factors to blocked format.
Returns list of L blocked tensors.
"""
rows, cols, L = sf.shape
n_row_blocks = ceil_div(rows, 128)
n_col_blocks = ceil_div(cols, 4)
# Single vectorized operation for all L
blocked = (sf.view(n_row_blocks, 128, n_col_blocks, 4, L)
.permute(0, 2, 1, 3, 4)
.reshape(-1, 4, 32, 4, L)
.transpose(1, 2)
.reshape(-1, 32, 16, L)
.flatten(0, 2)) # (blocked_size, L)
# Return list of contiguous slices
return [blocked[:, i].contiguous() for i in range(L)]
@torch.no_grad()
def to_blocked_single(sf: torch.Tensor) -> torch.Tensor:
"""Convert (rows, cols) scale factors to blocked format."""
rows, cols = sf.shape
n_row_blocks = ceil_div(rows, 128)
n_col_blocks = ceil_div(cols, 4)
return (sf.view(n_row_blocks, 128, n_col_blocks, 4)
.permute(0, 2, 1, 3)
.reshape(-1, 4, 32, 4)
.transpose(1, 2)
.reshape(-1, 32, 16)
.flatten())
@torch.no_grad()
def custom_kernel(data: input_t) -> output_t:
"""Simple optimized NVFP4 block-scaled GEMV."""
a, b, sfa, sfb, _, _, c = data
_, _, l = c.shape
if l == 1:
# Single batch path
scale_a = to_blocked_single(sfa[:, :, 0])
scale_b = to_blocked_single(sfb[:, :, 0])
result = torch._scaled_mm(
a[:, :, 0], b[:, :, 0].t(),
scale_a, scale_b,
bias=None, out_dtype=torch.float16
)
c[:, 0, 0] = result[:, 0]
else:
# Multi-batch path with vectorized scale conversion
blocked_sfa = to_blocked_all(sfa)
blocked_sfb = to_blocked_all(sfb)
for i in range(l):
result = torch._scaled_mm(
a[:, :, i], b[:, :, i].t(),
blocked_sfa[i], blocked_sfb[i],
bias=None, out_dtype=torch.float16
)
c[:, 0, i] = result[:, 0]
return cscrolls · 78 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON