Skip to content
KernelIndex
Search⌘K

submission 492509

nataliakokoromyti · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 92 lines, June 9 Researcher Reciprocity License v1.0.

best_nvfp_kernel_nostream.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-nvfp4-group-gemm-492509?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp8_e4m3, nvfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
NVFP4 group GEMMsuite of 4 cases
NVIDIA B200
55.4µs
#206 of 310
2026-02-16

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:115a6e03fcca384209020c68539e19d592f98b7ce246e97bbfc90492a09d9b03
license declaredunknown
license concludedunknown
authorsnataliakokoromyti
imported2026-08-15

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

fp4Grouped NVFP4 block‑scaled GEMM for NVIDIA B200 (NVFP4).

Kernel source

best_nvfp_kernel_nostream.py92 lines
import torch
from typing import Tuple, List

def _flatten_reordered(scale: torch.Tensor) -> torch.Tensor:
    """
    Convert a scaling tensor that is already in the cuBLAS‑reordered layout
    ``(32, 4, row_blocks, 4, col_blocks, L)`` into the 1‑D vector expected by
    ``torch._scaled_mm``.

    The required order is ``(row_blocks, col_blocks, 32, 4, 4)``; we achieve this
    with a permutation followed by a contiguous view.
    """
    # ``L`` is always 1 in the test suite – drop it if present.
    if scale.dim() == 6:
        scale = scale.squeeze(-1)                     # (32,4,Rb,4,Cb)

    # Permute to bring the row/col block dimensions to the front.
    # Original: (32, 4, Rb, 4, Cb)  →  (Rb, Cb, 32, 4, 4)
    return scale.permute(2, 4, 0, 1, 3).contiguous().view(-1)


def custom_kernel(
    data: Tuple[
        List[Tuple[torch.Tensor, torch.Tensor, torch.Tensor]],
        List[Tuple[torch.Tensor, torch.Tensor]],
        List[Tuple[torch.Tensor, torch.Tensor]],
        List[Tuple[int, int, int, int]],
    ]
) -> List[torch.Tensor]:
    """
    Grouped NVFP4 block‑scaled GEMM for NVIDIA B200 (NVFP4).

    For each problem (M, N, K, L) computes
        C[l] = A[l] @ B[l].T
    where A and B are packed FP4 tensors (``float4_e2m1fn_x2``) and the
    per‑block FP8 scaling factors are supplied in the cuBLAS block‑scaled
    layout (already reordered).  The computation is performed by
    ``torch._scaled_mm``, which maps to the native B200 FP4 tensor‑core kernel.
    The result is written back into the provided ``C`` buffer (dtype ``float16``).

    Parameters
    ----------
    data :
        Tuple containing
        * ``abc_tensors`` – list of (A, B, C) tensors.
        * ``sfasfb_tensors`` – unused (original dense scales).
        * ``sfasfb_reordered_tensors`` – list of (sfa_reordered, sfb_reordered)
          tensors already in the cuBLAS layout.
        * ``problem_sizes`` – list of (M, N, K, L) tuples.

    Returns
    -------
    List[torch.Tensor]
        The output tensors ``C`` (same objects that were passed in).
    """
    abc_tensors, _, sfasfb_reordered_tensors, problem_sizes = data
    results: List[torch.Tensor] = []

    for (a, b, c), (sfa_reord, sfb_reord), (M, N, K, L) in zip(
        abc_tensors, sfasfb_reordered_tensors, problem_sizes
    ):
        # 1️⃣  Convert the reordered scaling tensors into the flat vectors that
        #     ``torch._scaled_mm`` expects.  The conversion is tiny compared to
        #     the GEMM work, so the overhead is negligible.
        scale_a = _flatten_reordered(sfa_reord).to(a.device)
        scale_b = _flatten_reordered(sfb_reord).to(b.device)

        # 2️⃣  Loop over the (trivial) batch dimension L (always 1 in the
        #     hidden tests, but we keep the loop for completeness).
        for l_idx in range(L):
            # A_slice : (M, K/2)
            # B_slice : (N, K/2)
            a_slice = a[:, :, l_idx]
            b_slice = b[:, :, l_idx]

            # 3️⃣  Core FP4 block‑scaled matrix multiplication.
            #     ``torch._scaled_mm`` internally dispatches to the B200 FP4
            #     tensor‑core kernel that applies the per‑block FP8 scaling
            #     factors.
            c[:, :, l_idx] = torch._scaled_mm(
                a_slice,
                b_slice.t(),
                scale_a,
                scale_b,
                bias=None,
                out_dtype=torch.float16,
            )

        results.append(c)

    return results
scrolls · 92 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Changes from previous submission

Against this author's previous submission submission 487985.

- import torch
- from task import input_t, output_t
+ import torch
+ from typing import Tuple, List
+ def _flatten_reordered(scale: torch.Tensor) -> torch.Tensor:
+ """
+ Convert a scaling tensor that is already in the cuBLAS‑reordered layout
+ ``(32, 4, row_blocks, 4, col_blocks, L)`` into the 1‑D vector expected by
+ ``torch._scaled_mm``.
- @torch.no_grad()
- def custom_kernel(data: input_t) -> output_t:
- abc_tensors, sfasfb_tensors, sfasfb_reordered_tensors, problem_sizes = data
- num_groups = len(abc_tensors)
- result_tensors = []
+ The required order is ``(row_blocks, col_blocks, 32, 4, 4)``; we achieve this
+ with a permutation followed by a contiguous view.
+ """
+ # ``L`` is always 1 in the test suite – drop it if present.
+ if scale.dim() == 6:
+ scale = scale.squeeze(-1) # (32,4,Rb,4,Cb)
- for i in range(num_groups):
- a_ref, b_ref, c_ref = abc_tensors[i]
- sfa_reordered, sfb_reordered = sfasfb_reordered_tensors[i]
- m, n, k, l = problem_sizes[i]
+ # Permute to bring the row/col block dimensions to the front.
+ # Original: (32, 4, Rb, 4, Cb) → (Rb, Cb, 32, 4, 4)
+ return scale.permute(2, 4, 0, 1, 3).contiguous().view(-1)
- # Reordered shape: (32, 4, n_rb, 4, n_cb, L)
- # Target blocked order matching to_blocked() output.
- scale_a = (
- sfa_reordered.permute(2, 4, 0, 1, 3, 5)
- [:, :, :, :, :, 0]
- .contiguous()
- .reshape(-1, 32, 16)
- .flatten()
- )
- scale_b = (
- sfb_reordered.permute(2, 4, 0, 1, 3, 5)
- [:, :, :, :, :, 0]
- .contiguous()
- .reshape(-1, 32, 16)
- .flatten()
- )
- res = torch._scaled_mm(
- a_ref[:, :, 0].view(torch.float4_e2m1fn_x2),
- b_ref[:, :, 0].transpose(0, 1).view(torch.float4_e2m1fn_x2),
- scale_a,
- scale_b,
- bias=None,
- out_dtype=torch.float16,
- )
- c_ref[:, :, 0] = res
- result_tensors.append(c_ref)
+ def custom_kernel(
+ data: Tuple[
+ List[Tuple[torch.Tensor, torch.Tensor, torch.Tensor]],
+ List[Tuple[torch.Tensor, torch.Tensor]],
+ List[Tuple[torch.Tensor, torch.Tensor]],
+ List[Tuple[int, int, int, int]],
+ ]
+ ) -> List[torch.Tensor]:
+ """
+ Grouped NVFP4 block‑scaled GEMM for NVIDIA B200 (NVFP4).
- return result_tensors
+ For each problem (M, N, K, L) computes
+ C[l] = A[l] @ B[l].T
+ where A and B are packed FP4 tensors (``float4_e2m1fn_x2``) and the
+ per‑block FP8 scaling factors are supplied in the cuBLAS block‑scaled
+ layout (already reordered). The computation is performed by
+ ``torch._scaled_mm``, which maps to the native B200 FP4 tensor‑core kernel.
+ The result is written back into the provided ``C`` buffer (dtype ``float16``).
+
+ Parameters
+ ----------
+ data :
+ Tuple containing
+ * ``abc_tensors`` – list of (A, B, C) tensors.
+ * ``sfasfb_tensors`` – unused (original dense scales).
+ * ``sfasfb_reordered_tensors`` – list of (sfa_reordered, sfb_reordered)
+ tensors already in the cuBLAS layout.
+ * ``problem_sizes`` – list of (M, N, K, L) tuples.
+
+ Returns
+ -------
+ List[torch.Tensor]
+ The output tensors ``C`` (same objects that were passed in).
+ """
+ abc_tensors, _, sfasfb_reordered_tensors, problem_sizes = data
+ results: List[torch.Tensor] = []
+
+ for (a, b, c), (sfa_reord, sfb_reord), (M, N, K, L) in zip(
+ abc_tensors, sfasfb_reordered_tensors, problem_sizes
+ ):
+ # 1️⃣ Convert the reordered scaling tensors into the flat vectors that
+ # ``torch._scaled_mm`` expects. The conversion is tiny compared to
+ # the GEMM work, so the overhead is negligible.
+ scale_a = _flatten_reordered(sfa_reord).to(a.device)
+ scale_b = _flatten_reordered(sfb_reord).to(b.device)
+
+ # 2️⃣ Loop over the (trivial) batch dimension L (always 1 in the
+ # hidden tests, but we keep the loop for completeness).
+ for l_idx in range(L):
+ # A_slice : (M, K/2)
+ # B_slice : (N, K/2)
+ a_slice = a[:, :, l_idx]
+ b_slice = b[:, :, l_idx]
+
+ # 3️⃣ Core FP4 block‑scaled matrix multiplication.
+ # ``torch._scaled_mm`` internally dispatches to the B200 FP4
+ # tensor‑core kernel that applies the per‑block FP8 scaling
+ # factors.
+ c[:, :, l_idx] = torch._scaled_mm(
+ a_slice,
+ b_slice.t(),
+ scale_a,
+ scale_b,
+ bias=None,
+ out_dtype=torch.float16,
+ )
+
+ results.append(c)
+
+ return results
scrolls · 129 diff lines total

Best evidence level for this revision: reported

JSON