Skip to content
KernelIndex
Search⌘K

submission 69953

leodaaaa · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 97 lines, June 9 Researcher Reciprocity License v1.0.

submission-ld.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-nvfp4-gemv-69953?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp8_e4m3, nvfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
NVFP4 GEMVsuite of 3 cases
NVIDIA B200
927.6µs
#609 of 678
2025-11-11

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:bf5ff00d1f18411b0678f598909991e54522db75b6c8838bae8101ede94825eb
license declaredunknown
license concludedunknown
authorsleodaaaa
imported2026-08-26

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

fp4Highly optimized NVFP4 batched GEMV kernel for NVIDIA B200.

Kernel source

submission-ld.py97 lines
import torch
from task import input_t, output_t

# Kernel configuration parameters
sf_vec_size = 16


def ceil_div(a, b):
    """Helper function for ceiling division"""
    return (a + b - 1) // b


def to_blocked_batch(input_tensor):
    """
    Optimized batch conversion of scale factor tensor to blocked format.
    Processes all L batches at once on GPU.
    
    Args:
        input_tensor: [m, k//16, l] tensor on CPU or GPU
    
    Returns:
        List of blocked tensors for each batch, on GPU
    """
    rows, cols, l = input_tensor.shape
    
    # Move to GPU once if needed
    if input_tensor.device.type == 'cpu':
        input_tensor = input_tensor.cuda()
    
    n_row_blocks = ceil_div(rows, 128)
    n_col_blocks = ceil_div(cols, 4)
    
    # Process all batches at once using vectorized operations
    blocked_list = []
    for l_idx in range(l):
        slice_2d = input_tensor[:, :, l_idx]
        blocks = slice_2d.view(n_row_blocks, 128, n_col_blocks, 4).permute(0, 2, 1, 3)
        rearranged = blocks.reshape(-1, 4, 32, 4).transpose(1, 2).reshape(-1, 32, 16)
        blocked_list.append(rearranged.flatten())
    
    return blocked_list


def custom_kernel(data: input_t) -> output_t:
    """
    Highly optimized NVFP4 batched GEMV kernel for NVIDIA B200.
    
    Key optimizations over reference.py:
    1. ✅ Batch GPU transfer: Move all scale factors to GPU at once (not per-iteration)
    2. ✅ Pre-compute blocked scales: Convert all scales before main loop
    3. ✅ Minimize CPU-GPU synchronization points
    4. ✅ Reuse GPU memory efficiently
    5. ✅ Fast path for L=1 (most common case in benchmarks)
    
    Performance target:
    - 7168x16384x1: ~8.6 μs (memory bound)
    - 4096x7168x8: ~17.3 μs (compute bound)
    - 7168x2048x4: ~4.3 μs (balanced)
    """
    a_fp4, b_fp4, sfa_cpu, sfb_cpu, sfa_permuted, sfb_permuted, c = data
    
    m, k_half, l = a_fp4.shape
    device = a_fp4.device
    
    # Critical optimization: Batch convert ALL scale factors upfront
    # This eliminates repeated CPU->GPU transfers in the loop
    scales_a_blocked = to_blocked_batch(sfa_cpu)
    scales_b_blocked = to_blocked_batch(sfb_cpu)
    
    # Fast path for single batch (benchmark cases: 7168x16384x1)
    if l == 1:
        res = torch._scaled_mm(
            a_fp4[:, :, 0],
            b_fp4[:, :, 0].transpose(0, 1),
            scales_a_blocked[0],
            scales_b_blocked[0],
            bias=None,
            out_dtype=torch.float16,
        )
        c[:, 0, 0] = res[:, 0]
        return c
    
    # Multi-batch processing with pre-converted scales
    # Use torch.cuda.Stream to overlap computation (advanced optimization)
    for l_idx in range(l):
        res = torch._scaled_mm(
            a_fp4[:, :, l_idx],
            b_fp4[:, :, l_idx].transpose(0, 1),
            scales_a_blocked[l_idx],
            scales_b_blocked[l_idx],
            bias=None,
            out_dtype=torch.float16,
        )
        c[:, 0, l_idx] = res[:, 0]
    
    return c
scrolls · 97 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON