Skip to content
KernelIndex
Search⌘K

submission 105294

fl4res · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 78 lines, June 9 Researcher Reciprocity License v1.0.

fp4_nogcc.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-nvfp4-gemv-105294?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp8_e4m3, nvfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
NVFP4 GEMVsuite of 3 cases
NVIDIA B200
144.2µs
#484 of 678
2025-11-26

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:491b44f8c2502727086010a794a3c37ab5959b6f884b792619510680f7565f0d
license declaredunknown
license concludedunknown
authorsfl4res
imported2026-08-26

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

fp4Simple optimized NVFP4 block-scaled GEMV kernel.

Kernel source

fp4_nogcc.py78 lines
"""
Simple optimized NVFP4 block-scaled GEMV kernel.
No CUDA graphs - just pure vectorized operations.
"""
import torch
from task import input_t, output_t


def ceil_div(a: int, b: int) -> int:
    return (a + b - 1) // b


@torch.no_grad()
def to_blocked_all(sf: torch.Tensor) -> list:
    """
    Convert (rows, cols, L) scale factors to blocked format.
    Returns list of L blocked tensors.
    """
    rows, cols, L = sf.shape
    n_row_blocks = ceil_div(rows, 128)
    n_col_blocks = ceil_div(cols, 4)
    
    # Single vectorized operation for all L
    blocked = (sf.view(n_row_blocks, 128, n_col_blocks, 4, L)
                 .permute(0, 2, 1, 3, 4)
                 .reshape(-1, 4, 32, 4, L)
                 .transpose(1, 2)
                 .reshape(-1, 32, 16, L)
                 .flatten(0, 2))  # (blocked_size, L)
    
    # Return list of contiguous slices
    return [blocked[:, i].contiguous() for i in range(L)]


@torch.no_grad()
def to_blocked_single(sf: torch.Tensor) -> torch.Tensor:
    """Convert (rows, cols) scale factors to blocked format."""
    rows, cols = sf.shape
    n_row_blocks = ceil_div(rows, 128)
    n_col_blocks = ceil_div(cols, 4)
    return (sf.view(n_row_blocks, 128, n_col_blocks, 4)
              .permute(0, 2, 1, 3)
              .reshape(-1, 4, 32, 4)
              .transpose(1, 2)
              .reshape(-1, 32, 16)
              .flatten())


@torch.no_grad()
def custom_kernel(data: input_t) -> output_t:
    """Simple optimized NVFP4 block-scaled GEMV."""
    a, b, sfa, sfb, _, _, c = data
    _, _, l = c.shape
    
    if l == 1:
        # Single batch path
        scale_a = to_blocked_single(sfa[:, :, 0])
        scale_b = to_blocked_single(sfb[:, :, 0])
        result = torch._scaled_mm(
            a[:, :, 0], b[:, :, 0].t(),
            scale_a, scale_b,
            bias=None, out_dtype=torch.float16
        )
        c[:, 0, 0] = result[:, 0]
    else:
        # Multi-batch path with vectorized scale conversion
        blocked_sfa = to_blocked_all(sfa)
        blocked_sfb = to_blocked_all(sfb)
        
        for i in range(l):
            result = torch._scaled_mm(
                a[:, :, i], b[:, :, i].t(),
                blocked_sfa[i], blocked_sfb[i],
                bias=None, out_dtype=torch.float16
            )
            c[:, 0, i] = result[:, 0]
    
    return c
scrolls · 78 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON