Skip to content
KernelIndex
Search⌘K

submission 526982

ennkod · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 40 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-mxfp4-mm-526982?include=source"
interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, mxfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
AMD MXFP4 GEMMsuite of 6 cases
AMD Instinct MI355X
15.0µs
#568 of 1143
2026-03-10

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:d5b35839d3b68095958e2f1100ab8d38c97650ce05c29e05d89db883ca60bb02
license declaredunknown
license concludedunknown
authorsennkod
imported2026-08-26

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

fp4MXFP4 matrix multiplication on AMD MI355X.

Kernel source

submission.py40 lines
#!POPCORN leaderboard amd-mxfp4-mm
#!POPCORN gpu MI355X

"""
MXFP4 matrix multiplication on AMD MI355X.

Pipeline:
  bf16 A  ->  per-1x32 MXFP4 quant (with hardware shuffle)  ->  a4w4 GEMM  ->  bf16 C

B is provided pre-quantized and pre-shuffled (B_shuffle, B_scale_sh), so only A
needs to be quantized at runtime.  The quantization function is compiled once at
module import time to avoid Triton JIT overhead on the first timed call.
"""

import aiter
from aiter import QuantType, dtypes
from task import input_t, output_t

# Compile the Triton quantization kernel once at import time.
# Keeps it out of the timed hot path and ensures the JIT cache is warm.
_quant_func = aiter.get_triton_quant(QuantType.per_1x32)


def custom_kernel(data: input_t) -> output_t:
    A, B, B_q, B_shuffle, B_scale_sh = data

    # Quantize A: bf16 -> MXFP4 (e2m1) with per-1x32 scales.
    # shuffle=True produces the layout expected by gemm_a4w4 with bpreshuffle=True.
    A_q, A_scale_sh = _quant_func(A.contiguous(), shuffle=True)

    # Hardware-accelerated 4-bit GEMM using pre-shuffled B weights.
    return aiter.gemm_a4w4(
        A_q,
        B_shuffle,
        A_scale_sh,
        B_scale_sh,
        dtype=dtypes.bf16,
        bpreshuffle=True,
    )
scrolls · 40 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON