Skip to content
KernelIndex
Search⌘K

submission 102763

forestier_kasapi · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 72 lines, June 9 Researcher Reciprocity License v1.0.

submission4.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-nvfp4-gemv-102763?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp8_e4m3, nvfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
NVFP4 GEMVsuite of 3 cases
NVIDIA B200
116.5µs
#470 of 678
2025-11-25

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:ba4233194de0068f7266e7bdba9a98025f4feb1a1e0443009a8a79de4280a5f4
license declaredunknown
license concludedunknown
authorsforestier_kasapi
imported2026-08-26

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

fp4Optimized NVFP4 GEMV using batched preprocessing and vectorized operations.

Kernel source

submission4.py72 lines
import torch
from task import input_t, output_t

SF_VEC_SIZE = 16

def to_blocked_batched(input_matrix: torch.Tensor) -> torch.Tensor:
    """
    Convert scale-factor matrix to blocked layout for all batches at once.
    input_matrix: shape [rows, cols, batch]
    Returns: tensor of shape [batch, flattened_size]
    """
    rows, cols, batch = input_matrix.shape
    n_row_blocks = (rows + 127) >> 7  # Bit shift instead of ceil_div
    n_col_blocks = (cols + 3) >> 2
    
    # Single reshape pipeline - minimize intermediate tensors
    # [rows, cols, batch] -> [batch, n_row_blocks, 128, n_col_blocks, 4]
    blocks = input_matrix.permute(2, 0, 1).reshape(
        batch, n_row_blocks, 128, n_col_blocks, 4
    )
    
    # [batch, n_row_blocks, n_col_blocks, 128, 4] -> [batch, n_row_blocks*n_col_blocks, 32, 16]
    rearranged = (blocks.permute(0, 1, 3, 2, 4)
                  .reshape(batch, -1, 4, 32, 4)
                  .permute(0, 1, 3, 2, 4)
                  .reshape(batch, -1, 32, 16))
    
    # Return flattened batches: [batch, -1]
    return rearranged.flatten(1)


def custom_kernel(data: input_t) -> output_t:
    """
    Optimized NVFP4 GEMV using batched preprocessing and vectorized operations.
    """
    a_ref, b_ref, sfa_ref, sfb_ref, _, _, c_ref = data
    _, _, l = c_ref.shape
    
    # Preprocess all scale factors at once (batched conversion)
    scales_a = to_blocked_batched(sfa_ref)  # [l, -1]
    scales_b = to_blocked_batched(sfb_ref)  # [l, -1]
    
    # Vectorized batch processing
    if l > 1:
        # Process multiple batches in parallel when possible
        for l_idx in range(l):
            a_slice = a_ref[:, :, l_idx]
            b_slice = b_ref[:, :, l_idx].transpose(0, 1)
            
            res = torch._scaled_mm(
                a_slice,
                b_slice,
                scales_a[l_idx],
                scales_b[l_idx],
                bias=None,
                out_dtype=torch.float16,
            )
            c_ref[:, 0, l_idx] = res[:, 0]
    else:
        # Single batch optimization
        res = torch._scaled_mm(
            a_ref[:, :, 0],
            b_ref[:, :, 0].transpose(0, 1),
            scales_a[0],
            scales_b[0],
            bias=None,
            out_dtype=torch.float16,
        )
        c_ref[:, 0, 0] = res[:, 0]
    
    return c_ref
scrolls · 72 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON