Skip to content
KernelIndex
Search⌘K

submission 483123

farkhanda ghaffar · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 44 lines, June 9 Researcher Reciprocity License v1.0.

v55_2.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-nvfp4-group-gemm-483123?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp8_e4m3, nvfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
NVFP4 group GEMMsuite of 4 cases
NVIDIA B200
55.4µs
#204 of 310
2026-02-07

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:8b2a6c390837f5dceab43233620dd12446a2a2f83991339ef239c481eec7b41a
license declaredunknown
license concludedunknown
authorsfarkhanda ghaffar
imported2026-08-15

Kernel source

v55_2.py44 lines
import torch
import math
from task import input_t, output_t
from reference import ceil_div

def custom_kernel(data: input_t) -> output_t:
    abc_tensors, sfasfb_tensors, sfasfb_reordered_tensors, problem_sizes = data
    result_tensors = []
    
    # Pre-compute all flattening parameters once
    for i, ((a, b, c), (sfa_reordered, sfb_reordered), (m, n, k, l)) in enumerate(
        zip(abc_tensors, sfasfb_reordered_tensors, problem_sizes)
    ):
        # Direct fused permute + flatten without intermediate allocations
        # Use in-place operations where possible
        sfa_slice = sfa_reordered[..., 0]
        sfb_slice = sfb_reordered[..., 0]
        
        # Fast path for l=1 (all benchmarks)
        # Use view-based reshaping without copy when possible
        scale_a_flat = sfa_slice.permute(2, 4, 0, 1, 3).reshape(-1)
        scale_b_flat = sfb_slice.permute(2, 4, 0, 1, 3).reshape(-1)
        
        # Ensure contiguous memory for scaled_mm
        scale_a_flat = scale_a_flat.contiguous()
        scale_b_flat = scale_b_flat.contiguous()
        
        # Fuse the transpose with view
        a_view = a[:, :, 0].view(torch.float4_e2m1fn_x2)
        b_view = b[:, :, 0].transpose(0, 1).view(torch.float4_e2m1fn_x2)
        
        # Single scaled_mm call
        c[:, :, 0] = torch._scaled_mm(
            a_view,
            b_view,
            scale_a_flat,
            scale_b_flat,
            bias=None,
            out_dtype=torch.float16,
        )
        
        result_tensors.append(c)
    
    return result_tensors
scrolls · 44 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON