Skip to content
KernelIndex
Search⌘K

submission 514005

ohamnl. · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 66 lines, June 9 Researcher Reciprocity License v1.0.

solution.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-mxfp4-mm-514005?include=source"
interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, mxfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
AMD MXFP4 GEMMsuite of 6 cases
AMD Instinct MI355X
15.2µs
#610 of 1143
2026-03-07

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:7edef0b75e522cd0016215287054e63665ddb7f5d19d8a22351594db6e4f1ba2
license declaredunknown
license concludedunknown
authorsohamnl.
imported2026-08-26

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

fp4MXFP4 GEMM: bf16 A x MXFP4 B -> bf16 C

Kernel source

solution.py66 lines
#!POPCORN leaderboard amd-mxfp4-mm
#!POPCORN gpu MI355X
import torch
from task import input_t, output_t
from utils import make_match_reference

import aiter
from aiter import QuantType, dtypes
from aiter.ops.shuffle import shuffle_weight

# the ref quantizes A fresh every call with get_triton_quant
# that's the only real cost here since B is already pre-shuffled
# our job: make A quant + gemm as fast as possible

# cache the quant function — get_triton_quant does jit compilation on first call
# calling it once at module load means zero compilation overhead at runtime
_quant_func = None

def _get_quant():
    global _quant_func
    if _quant_func is None:
        _quant_func = aiter.get_triton_quant(QuantType.per_1x32)
    return _quant_func


def custom_kernel(data: input_t) -> output_t:
    """
    MXFP4 GEMM: bf16 A x MXFP4 B -> bf16 C

    ref_kernel quantizes A with get_triton_quant every single call.
    We cache the quant function to avoid recompilation overhead.
    B is already shuffled in generate_input — use it directly.

    The real kernel is aiter.gemm_a4w4 which maps to CDNA4 native
    MXFP4 tensor core instructions. We can't replace that.
    What we can do: make sure everything feeding into it is optimal.
    """
    A, B, B_q, B_shuffle, B_scale_sh = data

    # contiguous check — gemm_a4w4 requires contiguous input
    # in the ref this is called unconditionally, we only do it if needed
    if not A.is_contiguous():
        A = A.contiguous()

    # quantize A to MXFP4 with shuffled layout
    # shuffle=True matches what gemm_a4w4 expects with bpreshuffle=True
    A_q, A_scale_sh = _get_quant()(A, shuffle=True)

    # B_shuffle and B_scale_sh are pre-computed in generate_input
    # zero extra work here vs ref which also uses them directly
    out = aiter.gemm_a4w4(
        A_q,
        B_shuffle,
        A_scale_sh,
        B_scale_sh,
        dtype=dtypes.bf16,
        bpreshuffle=True,
    )

    return out


solution = custom_kernel

from reference import ref_kernel
check_implementation = make_match_reference(ref_kernel, rtol=1e-02, atol=1e-02)
scrolls · 66 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON