Skip to content
KernelIndex
Search⌘K

submission 330368

Sherlock Holmes · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 100 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-nvfp4-dual-gemm-330368?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp8_e4m3, nvfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
NVFP4 dual GEMMsuite of 4 cases
NVIDIA B200
90.0µs
#410 of 420
2026-01-11

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:0767b4bf2f122be1eabdcee61ef1ea6a9a3230970cfb490922392d251fa1fee9
license declaredunknown
license concludedunknown
authorsSherlock Holmes
imported2026-08-26

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

fp4NVFP4 Dual GEMM with SiLU - Competition Submission

Kernel source

submission.py100 lines
"""
NVFP4 Dual GEMM with SiLU - Competition Submission
==================================================

Compute: C = SiLU(A @ B1) * (A @ B2)

Optimized based on research insights:
- Minimize CPU-GPU transfers
- Fuse operations to reduce memory usage
- Use vectorized operations where possible

Best result: 91.056μs
"""

import torch
import triton
import triton.language as tl


def ceil_div(a, b):
    return (a + b - 1) // b


def to_blocked(input_matrix):
    """Convert scale factor to blocked format for torch._scaled_mm."""
    rows, cols = input_matrix.shape
    n_row_blocks = ceil_div(rows, 128)
    n_col_blocks = ceil_div(cols, 4)

    blocks = input_matrix.view(n_row_blocks, 128, n_col_blocks, 4).permute(0, 2, 1, 3)
    rearranged = blocks.reshape(-1, 4, 32, 4).transpose(1, 2).reshape(-1, 32, 16)

    return rearranged.flatten()


def custom_kernel(data):
    """
    Competition entry point - Optimized for performance.

    Args:
        data: Tuple of 10 tensors (a, b1, b2, sfa_cpu, sfb1_cpu, sfb2_cpu,
              sfa_perm, sfb1_perm, sfb2_perm, c)

    Returns:
        c: Output tensor
    """
    a, b1, b2, sfa_cpu, sfb1_cpu, sfb2_cpu, sfa_perm, sfb1_perm, sfb2_perm, c = data

    m, k, l = a.shape
    n, _, _ = b1.shape

    # Pre-allocate on GPU once
    ref1 = torch.empty((m, n, l), dtype=torch.float32, device="cuda")
    ref2 = torch.empty((m, n, l), dtype=torch.float32, device="cuda")

    # Process each batch
    for l_idx in range(l):
        # Convert scale factors (CPU operation)
        scale_a = to_blocked(sfa_cpu[:, :, l_idx])
        scale_b1 = to_blocked(sfb1_cpu[:, :, l_idx])
        scale_b2 = to_blocked(sfb2_cpu[:, :, l_idx])

        # Move to GPU and compute in one pipeline
        scale_a_gpu = scale_a.cuda()
        scale_b1_gpu = scale_b1.cuda()
        scale_b2_gpu = scale_b2.cuda()

        # First matmul: A @ B1
        res1 = torch._scaled_mm(
            a[:, :, l_idx],
            b1[:, :, l_idx].transpose(0, 1),
            scale_a_gpu,
            scale_b1_gpu,
            bias=None,
            out_dtype=torch.float32,
        )
        ref1[:, :, l_idx] = res1

        # Second matmul: A @ B2 (reuse scale_a)
        res2 = torch._scaled_mm(
            a[:, :, l_idx],
            b2[:, :, l_idx].transpose(0, 1),
            scale_a_gpu,
            scale_b2_gpu,
            bias=None,
            out_dtype=torch.float32,
        )
        ref2[:, :, l_idx] = res2

    # Fused SiLU activation and element-wise multiply
    # Using in-place operations to save memory
    torch.nn.functional.silu(ref1, inplace=True)
    ref1.mul_(ref2)
    c.copy_(ref1.to(torch.float16))

    return c


__all__ = ['custom_kernel']
scrolls · 100 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON