submission 492509
nataliakokoromyti · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 92 lines, June 9 Researcher Reciprocity License v1.0.
best_nvfp_kernel_nostream.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-nvfp4-group-gemm-492509?include=source"interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp8_e4m3, nvfp4
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:115a6e03fcca384209020c68539e19d592f98b7ce246e97bbfc90492a09d9b03
license declaredunknown
license concludedunknown
authorsnataliakokoromyti
imported2026-08-15
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
fp4
Grouped NVFP4 block‑scaled GEMM for NVIDIA B200 (NVFP4).Kernel source
best_nvfp_kernel_nostream.py92 lines
import torch
from typing import Tuple, List
def _flatten_reordered(scale: torch.Tensor) -> torch.Tensor:
"""
Convert a scaling tensor that is already in the cuBLAS‑reordered layout
``(32, 4, row_blocks, 4, col_blocks, L)`` into the 1‑D vector expected by
``torch._scaled_mm``.
The required order is ``(row_blocks, col_blocks, 32, 4, 4)``; we achieve this
with a permutation followed by a contiguous view.
"""
# ``L`` is always 1 in the test suite – drop it if present.
if scale.dim() == 6:
scale = scale.squeeze(-1) # (32,4,Rb,4,Cb)
# Permute to bring the row/col block dimensions to the front.
# Original: (32, 4, Rb, 4, Cb) → (Rb, Cb, 32, 4, 4)
return scale.permute(2, 4, 0, 1, 3).contiguous().view(-1)
def custom_kernel(
data: Tuple[
List[Tuple[torch.Tensor, torch.Tensor, torch.Tensor]],
List[Tuple[torch.Tensor, torch.Tensor]],
List[Tuple[torch.Tensor, torch.Tensor]],
List[Tuple[int, int, int, int]],
]
) -> List[torch.Tensor]:
"""
Grouped NVFP4 block‑scaled GEMM for NVIDIA B200 (NVFP4).
For each problem (M, N, K, L) computes
C[l] = A[l] @ B[l].T
where A and B are packed FP4 tensors (``float4_e2m1fn_x2``) and the
per‑block FP8 scaling factors are supplied in the cuBLAS block‑scaled
layout (already reordered). The computation is performed by
``torch._scaled_mm``, which maps to the native B200 FP4 tensor‑core kernel.
The result is written back into the provided ``C`` buffer (dtype ``float16``).
Parameters
----------
data :
Tuple containing
* ``abc_tensors`` – list of (A, B, C) tensors.
* ``sfasfb_tensors`` – unused (original dense scales).
* ``sfasfb_reordered_tensors`` – list of (sfa_reordered, sfb_reordered)
tensors already in the cuBLAS layout.
* ``problem_sizes`` – list of (M, N, K, L) tuples.
Returns
-------
List[torch.Tensor]
The output tensors ``C`` (same objects that were passed in).
"""
abc_tensors, _, sfasfb_reordered_tensors, problem_sizes = data
results: List[torch.Tensor] = []
for (a, b, c), (sfa_reord, sfb_reord), (M, N, K, L) in zip(
abc_tensors, sfasfb_reordered_tensors, problem_sizes
):
# 1️⃣ Convert the reordered scaling tensors into the flat vectors that
# ``torch._scaled_mm`` expects. The conversion is tiny compared to
# the GEMM work, so the overhead is negligible.
scale_a = _flatten_reordered(sfa_reord).to(a.device)
scale_b = _flatten_reordered(sfb_reord).to(b.device)
# 2️⃣ Loop over the (trivial) batch dimension L (always 1 in the
# hidden tests, but we keep the loop for completeness).
for l_idx in range(L):
# A_slice : (M, K/2)
# B_slice : (N, K/2)
a_slice = a[:, :, l_idx]
b_slice = b[:, :, l_idx]
# 3️⃣ Core FP4 block‑scaled matrix multiplication.
# ``torch._scaled_mm`` internally dispatches to the B200 FP4
# tensor‑core kernel that applies the per‑block FP8 scaling
# factors.
c[:, :, l_idx] = torch._scaled_mm(
a_slice,
b_slice.t(),
scale_a,
scale_b,
bias=None,
out_dtype=torch.float16,
)
results.append(c)
return results
scrolls · 92 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Changes from previous submission
Against this author's previous submission submission 487985.
- import torch- from task import input_t, output_t+ import torch+ from typing import Tuple, List+ def _flatten_reordered(scale: torch.Tensor) -> torch.Tensor:+ """+ Convert a scaling tensor that is already in the cuBLAS‑reordered layout+ ``(32, 4, row_blocks, 4, col_blocks, L)`` into the 1‑D vector expected by+ ``torch._scaled_mm``.- @torch.no_grad()- def custom_kernel(data: input_t) -> output_t:- abc_tensors, sfasfb_tensors, sfasfb_reordered_tensors, problem_sizes = data- num_groups = len(abc_tensors)- result_tensors = []+ The required order is ``(row_blocks, col_blocks, 32, 4, 4)``; we achieve this+ with a permutation followed by a contiguous view.+ """+ # ``L`` is always 1 in the test suite – drop it if present.+ if scale.dim() == 6:+ scale = scale.squeeze(-1) # (32,4,Rb,4,Cb)- for i in range(num_groups):- a_ref, b_ref, c_ref = abc_tensors[i]- sfa_reordered, sfb_reordered = sfasfb_reordered_tensors[i]- m, n, k, l = problem_sizes[i]+ # Permute to bring the row/col block dimensions to the front.+ # Original: (32, 4, Rb, 4, Cb) → (Rb, Cb, 32, 4, 4)+ return scale.permute(2, 4, 0, 1, 3).contiguous().view(-1)- # Reordered shape: (32, 4, n_rb, 4, n_cb, L)- # Target blocked order matching to_blocked() output.- scale_a = (- sfa_reordered.permute(2, 4, 0, 1, 3, 5)- [:, :, :, :, :, 0]- .contiguous()- .reshape(-1, 32, 16)- .flatten()- )- scale_b = (- sfb_reordered.permute(2, 4, 0, 1, 3, 5)- [:, :, :, :, :, 0]- .contiguous()- .reshape(-1, 32, 16)- .flatten()- )- res = torch._scaled_mm(- a_ref[:, :, 0].view(torch.float4_e2m1fn_x2),- b_ref[:, :, 0].transpose(0, 1).view(torch.float4_e2m1fn_x2),- scale_a,- scale_b,- bias=None,- out_dtype=torch.float16,- )- c_ref[:, :, 0] = res- result_tensors.append(c_ref)+ def custom_kernel(+ data: Tuple[+ List[Tuple[torch.Tensor, torch.Tensor, torch.Tensor]],+ List[Tuple[torch.Tensor, torch.Tensor]],+ List[Tuple[torch.Tensor, torch.Tensor]],+ List[Tuple[int, int, int, int]],+ ]+ ) -> List[torch.Tensor]:+ """+ Grouped NVFP4 block‑scaled GEMM for NVIDIA B200 (NVFP4).- return result_tensors+ For each problem (M, N, K, L) computes+ C[l] = A[l] @ B[l].T+ where A and B are packed FP4 tensors (``float4_e2m1fn_x2``) and the+ per‑block FP8 scaling factors are supplied in the cuBLAS block‑scaled+ layout (already reordered). The computation is performed by+ ``torch._scaled_mm``, which maps to the native B200 FP4 tensor‑core kernel.+ The result is written back into the provided ``C`` buffer (dtype ``float16``).++ Parameters+ ----------+ data :+ Tuple containing+ * ``abc_tensors`` – list of (A, B, C) tensors.+ * ``sfasfb_tensors`` – unused (original dense scales).+ * ``sfasfb_reordered_tensors`` – list of (sfa_reordered, sfb_reordered)+ tensors already in the cuBLAS layout.+ * ``problem_sizes`` – list of (M, N, K, L) tuples.++ Returns+ -------+ List[torch.Tensor]+ The output tensors ``C`` (same objects that were passed in).+ """+ abc_tensors, _, sfasfb_reordered_tensors, problem_sizes = data+ results: List[torch.Tensor] = []++ for (a, b, c), (sfa_reord, sfb_reord), (M, N, K, L) in zip(+ abc_tensors, sfasfb_reordered_tensors, problem_sizes+ ):+ # 1️⃣ Convert the reordered scaling tensors into the flat vectors that+ # ``torch._scaled_mm`` expects. The conversion is tiny compared to+ # the GEMM work, so the overhead is negligible.+ scale_a = _flatten_reordered(sfa_reord).to(a.device)+ scale_b = _flatten_reordered(sfb_reord).to(b.device)++ # 2️⃣ Loop over the (trivial) batch dimension L (always 1 in the+ # hidden tests, but we keep the loop for completeness).+ for l_idx in range(L):+ # A_slice : (M, K/2)+ # B_slice : (N, K/2)+ a_slice = a[:, :, l_idx]+ b_slice = b[:, :, l_idx]++ # 3️⃣ Core FP4 block‑scaled matrix multiplication.+ # ``torch._scaled_mm`` internally dispatches to the B200 FP4+ # tensor‑core kernel that applies the per‑block FP8 scaling+ # factors.+ c[:, :, l_idx] = torch._scaled_mm(+ a_slice,+ b_slice.t(),+ scale_a,+ scale_b,+ bias=None,+ out_dtype=torch.float16,+ )++ results.append(c)++ return results
scrolls · 129 diff lines total
Best evidence level for this revision: reported
JSON