Skip to content
KernelIndex
Search⌘K

submission 594929

Shlok · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 81 lines, June 9 Researcher Reciprocity License v1.0.

amd-mxfp4-mm.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-mxfp4-mm-594929?include=source"
interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, mxfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
AMD MXFP4 GEMMsuite of 6 cases
AMD Instinct MI355X
24.4µs
#1067 of 1143
2026-03-20

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:cb627ad330a0972d049e707678e1e420e6ac6227699e954abd45005152f228ce
license declaredunknown
license concludedunknown
authorsShlok
imported2026-08-26

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

fp4FP4 quant + FP4 GEMM reference: bf16 A, MXFP4 B -> MXFP4 per-1x32 quant A -> gemm_a4w4 -> bf16 C.

Kernel source

amd-mxfp4-mm.py81 lines
#!POPCORN leaderboard amd-mxfp4-mm

# This is a submission template for popcorn leaderboard 'amd-mxfp4-mm'.
# Your task is as follows:
# > You will implement a quantize func and block scaled MXFP4 matrix-matrix multiplication kernel optimized for AMD Instinct MI355X GPU.
# > To be explicit, you will be given a tuple of tensors:
# > ```
# > (A, B, B_q, B_shuffle, B_scale_sh)
# > ```
# > where:
# > * `A` is M x K in K-major order in bfloat16
# > * `B` is N x K in K-major order in bfloat16
# > * `B_q` is N x K/2 in K-major order in MXFP4
# > * `B_shuffle` is N x K/2 in shuffled order in MXFP4, shuffled to (16,16) tile coalesced
# > * `B_scale_sh` is * x K/32 in E8M0, * means it will be padded.
# > 
# > Then, the kernel flow is bf16 A, MXFP4 B -> MXFP4 per-1x32 quant A -> gemm_a4w4 -> BF16 C [m,n].
# > To be specific, the invocation flow is:
# > (1) Quant A to MXFP4: aiter.get_triton_quant(QuantType.per_1x32). 
# > (2) GEMM: aiter.gemm_a4w4.
# > m, n divisible by 64; k divisible by 64.
# > 
# > The ranking criteria is the geometric mean of the benchmark results.
# > Pls note that this is the elimination round, whoever rank top5 are selected into the next round, e2e optimization for deepseek-R1-MXFP4 and GPTOSS-MXFP4 mdoel
# > ```
# > The aiter performance is:
# > M   N    K   time[us]
# >   4 2880   512  8.198
# >  16 2112  7168 20.873
# >  32 4096   512  9.462
# >  32 2880   512  9.173
# >  64 7168  2048 12.738
# > 256 3072  1536 12.219
# > ```
# The deadline for this leaderboard is 2026-04-07 07:59:00+00:00

# You can automatically route this file to specific GPUs by adding a line
# `#!POPCORN gpus <GPUs>` to the header of this file.
# Happy hacking!

"""
FP4 quant + FP4 GEMM reference: bf16 A, MXFP4 B -> MXFP4 per-1x32 quant A -> gemm_a4w4 -> bf16 C.
Quant logic follows aiter op_tests/test_gemm_a4w4.py (get_triton_quant(QuantType.per_1x32)).
"""
from task import input_t, output_t


def custom_kernel(data: input_t) -> output_t:
    """
    Reference: MXFP4 per-1x32 quant on A; B_shuffle, B_scale_sh from generate_input.
    gemm_a4w4 with bpreshuffle=True.
    """
    import aiter
    from aiter import QuantType, dtypes
    from aiter.ops.triton.quant import dynamic_mxfp4_quant 
    from aiter.utility.fp4_utils import e8m0_shuffle

    def _quant_mxfp4(x, shuffle=True):
        x_fp4, bs_e8m0 = dynamic_mxfp4_quant(x)
        if shuffle:
            bs_e8m0 = e8m0_shuffle(bs_e8m0)
        return x_fp4.view(dtypes.fp4x2), bs_e8m0.view(dtypes.fp8_e8m0)
    
    A, B, B_q, B_shuffle, B_scale_sh = data
    A = A.contiguous()
    B = B.contiguous()
    m, k = A.shape
    n, _ = B.shape

    A_q, A_scale_sh = _quant_mxfp4(A, shuffle=True)
    out_gemm = aiter.gemm_a4w4(
        A_q,
        B_shuffle,
        A_scale_sh,
        B_scale_sh,
        dtype=dtypes.bf16,
        bpreshuffle=True,
    )
    return out_gemm

scrolls · 81 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON