Skip to content
KernelIndex
Search⌘K

submission 69462

Max Brashear · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 84 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-nvfp4-gemv-69462?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp8_e4m3, nvfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
NVFP4 GEMVsuite of 3 cases
NVIDIA B200
2.24ms
#667 of 678
2025-11-11

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:870e19fb056a416d2962d9830ca529cf6e202cb4808a70220de76dd336ed79d1
license declaredunknown
license concludedunknown
authorsMax Brashear
imported2026-08-26

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

fp4Batched NVFP4(e2m1) GEMV with block FP8 scales.

Kernel source

submission.py84 lines
import torch
from task import input_t, output_t

# NVTX for nicer traces during profiling (optional if unavailable)
try:
    from torch.cuda.nvtx import range as nvtx_range
except Exception:  # pragma: no cover
    class nvtx_range:  # fallback no-op
        def __init__(self, *_args, **_kwargs): pass
        def __enter__(self): return self
        def __exit__(self, *exc): return False


# ---- helpers (mirrors the reference layout conversion) ---------------------

_SF_VEC = 16  # block size in K used by scaling factors


def _ceil_div(a: int, b: int) -> int:
    return (a + b - 1) // b


def _to_blocked(input_matrix: torch.Tensor) -> torch.Tensor:
    """
    Reorders a (rows, cols) FP8 scale matrix to the 32x4x...x4x... blocked
    format expected by torch._scaled_mm on Hopper/Blackwell parts.

    This is kept byte-for-byte compatible with the reference implementation.
    """
    rows, cols = input_matrix.shape
    n_row_blocks = _ceil_div(rows, 128)
    n_col_blocks = _ceil_div(cols, 4)

    padded = input_matrix
    blocks = padded.view(n_row_blocks, 128, n_col_blocks, 4).permute(0, 2, 1, 3)
    rearranged = blocks.reshape(-1, 4, 32, 4).transpose(1, 2).reshape(-1, 32, 16)
    return rearranged.flatten()


# ---- main kernel -----------------------------------------------------------

@torch.no_grad()
def custom_kernel(data: input_t) -> output_t:
    """
    Batched NVFP4(e2m1) GEMV with block FP8 scales.

    Expects a tuple:
      (a, b, sfa_ref_cpu, sfb_ref_cpu, _sfa_perm, _sfb_perm, c)
    where a:[M,K,L] and b:[1,K,L] are torch.float4_e2m1fn_x2 in K-major order,
    sfa/sfb are FP8(e4m3fn) scale tensors (reference layout on CPU),
    and c is [M,1,L] in FP16.

    For correctness and portability, this implementation mirrors the reference:
    for each batch slice l, convert the FP8 scales to the blocked format and
    call torch._scaled_mm, then write the single-column result into c.
    """
    a_ref, b_ref, sfa_ref_cpu, sfb_ref_cpu, _sfa_perm, _sfb_perm, c_ref = data

    M, _, L = a_ref.shape
    # Allocate the output (we overwrite every element)
    out = torch.empty_like(c_ref)

    with nvtx_range("batched_scaled_gemv"):
        for l_idx in range(L):
            with nvtx_range(f"slice_{l_idx}"):
                # Convert the per-slice scales to blocked format (GPU tensors)
                scale_a = _to_blocked(sfa_ref_cpu[:, :, l_idx]).cuda()
                scale_b = _to_blocked(sfb_ref_cpu[:, :, l_idx]).cuda()

                # (M,K) @ (1,K)^T => (M,1); use out_dtype fp16 as required
                # b is provided as [1,K] so transpose along K into [K,1]
                res = torch._scaled_mm(
                    a_ref[:, :, l_idx],
                    b_ref[:, :, l_idx].transpose(0, 1),
                    scale_a,
                    scale_b,
                    bias=None,
                    out_dtype=torch.float16,
                )
                # Write result vector into output tensor
                out[:, 0, l_idx] = res[:, 0]

    return out
scrolls · 84 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON