Skip to content
KernelIndex
Search⌘K

submission 75968

arsrivish26691 · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 107 lines, June 9 Researcher Reciprocity License v1.0.

nvfp4tsk1vg.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-nvfp4-gemv-75968?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp8_e4m3, nvfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
NVFP4 GEMVsuite of 3 cases
NVIDIA B200
126.4µs
#477 of 678
2025-11-14

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:16a582313332c2a5b26c98f7c1de98a826ed5d044df69d0cdcaa9c9a10617653
license declaredunknown
license concludedunknown
authorsarsrivish26691
imported2026-08-26

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

fp4NVFP4 block-scaled GEMV using torch._scaled_mm, with:

Kernel source

nvfp4tsk1vg.py107 lines
import torch
from task import input_t, output_t


def ceil_div(a: int, b: int) -> int:
    return (a + b - 1) // b


def to_blocked_all_gpu(sf: torch.Tensor) -> torch.Tensor:
    """
    Batched, GPU version of the provided `to_blocked`.

    Input:
        sf: [M, K_sf, L] (fp8) on CUDA, with:
            - M multiple of 128
            - K_sf multiple of 4

    Output:
        scales_all: [L, num_scales_per_matrix]
        where scales_all[l] == to_blocked(sf[:, :, l]) from the reference.
    """
    assert sf.dim() == 3, f"Expected 3D scale tensor, got {sf.shape}"
    M, K_sf, L = sf.shape

    assert M % 128 == 0, f"M={M} must be divisible by 128"
    assert K_sf % 4 == 0, f"K_sf={K_sf} must be divisible by 4"

    n_row_blocks = M // 128
    n_col_blocks = K_sf // 4

    # Start from [M, K_sf, L] on GPU
    # 1) [M, K_sf, L] -> [n_row_blocks, 128, n_col_blocks, 4, L]
    blocks = sf.view(n_row_blocks, 128, n_col_blocks, 4, L)

    # 2) [n_row_blocks, n_col_blocks, 128, 4, L]
    blocks = blocks.permute(0, 2, 1, 3, 4)

    # 3) merge row-block dims: [B, 4, 32, 4, L], B = n_row_blocks * n_col_blocks
    blocks = blocks.reshape(-1, 4, 32, 4, L)

    # 4) transpose 4 x 32 -> 32 x 4: [B, 32, 4, 4, L]
    blocks = blocks.transpose(1, 2)

    # 5) merge the two 4s -> 16: [B, 32, 16, L]
    blocks = blocks.reshape(-1, 32, 16, L)

    # 6) move L in front, flatten per-matrix: [L, B, 32, 16] -> [L, B*32*16]
    blocks = blocks.permute(3, 0, 1, 2).contiguous()  # [L, B, 32, 16]
    L_, B, _, _ = blocks.shape
    return blocks.view(L_, B * 32 * 16)  # [L, num_scales]


def custom_kernel(data: input_t) -> output_t:
    """
    NVFP4 block-scaled GEMV using torch._scaled_mm, with:
      - scale factors moved to GPU once
      - scale blocking/swizzling done on GPU for all L at once
      - inner loop over L only runs `_scaled_mm` and a copy
    """
    a_ref, b_ref, sfa_ref_cpu, sfb_ref_cpu, _, _, c_ref = data

    device = a_ref.device  # should be "cuda"
    M, _, L = c_ref.shape

    # ---- Move scale tensors to GPU once ----
    # sfa_ref_cpu : [M, K//16, L] on CPU → GPU
    # sfb_ref_cpu : [128, K//16, L] on CPU → GPU
    sfa_gpu = sfa_ref_cpu.to(device=device, non_blocking=True)
    sfb_gpu = sfb_ref_cpu.to(device=device, non_blocking=True)

    # ---- Precompute blocked scale vectors for ALL L on GPU ----
    # These are equivalent to:
    #   for l in range(L): to_blocked(sfa_ref_cpu[:, :, l])
    scale_a_all = to_blocked_all_gpu(sfa_gpu)  # [L, num_scales_a] on GPU
    scale_b_all = to_blocked_all_gpu(sfb_gpu)  # [L, num_scales_b] on GPU

    # ---- Main GEMV loop over batch L ----
    # We still need one `_scaled_mm` per l (batches are independent),
    # but now there is:
    #   - no CPU swizzle work
    #   - no per-l .cuda() calls for scale vectors
    for l_idx in range(L):
        # A_l: [M, K]
        a_l = a_ref[:, :, l_idx]

        # B_l: original is [128, K/2, L] in nvfp4; reference uses transpose
        # b_ref[:, :, l].T: [K, 128]
        b_l = b_ref[:, :, l_idx].transpose(0, 1)

        sa = scale_a_all[l_idx]  # [num_scales_a] on GPU
        sb = scale_b_all[l_idx]  # [num_scales_b] on GPU

        # Compute (M, K) @ (K, 128) with NVFP4 + block scales
        res = torch._scaled_mm(
            a_l,
            b_l,
            sa,
            sb,
            bias=None,
            out_dtype=torch.float16,
        )  # [M, 128]

        # c_ref: [M, 1, L] – GEMV, real N=1, we only keep the first column
        c_ref[:, 0, l_idx].copy_(res[:, 0])

    return c_ref
scrolls · 107 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON