submission 526982
ennkod · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 40 lines, June 9 Researcher Reciprocity License v1.0.
submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-mxfp4-mm-526982?include=source"interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, mxfp4
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:d5b35839d3b68095958e2f1100ab8d38c97650ce05c29e05d89db883ca60bb02
license declaredunknown
license concludedunknown
authorsennkod
imported2026-08-26
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
fp4
MXFP4 matrix multiplication on AMD MI355X.Kernel source
submission.py40 lines
#!POPCORN leaderboard amd-mxfp4-mm
#!POPCORN gpu MI355X
"""
MXFP4 matrix multiplication on AMD MI355X.
Pipeline:
bf16 A -> per-1x32 MXFP4 quant (with hardware shuffle) -> a4w4 GEMM -> bf16 C
B is provided pre-quantized and pre-shuffled (B_shuffle, B_scale_sh), so only A
needs to be quantized at runtime. The quantization function is compiled once at
module import time to avoid Triton JIT overhead on the first timed call.
"""
import aiter
from aiter import QuantType, dtypes
from task import input_t, output_t
# Compile the Triton quantization kernel once at import time.
# Keeps it out of the timed hot path and ensures the JIT cache is warm.
_quant_func = aiter.get_triton_quant(QuantType.per_1x32)
def custom_kernel(data: input_t) -> output_t:
A, B, B_q, B_shuffle, B_scale_sh = data
# Quantize A: bf16 -> MXFP4 (e2m1) with per-1x32 scales.
# shuffle=True produces the layout expected by gemm_a4w4 with bpreshuffle=True.
A_q, A_scale_sh = _quant_func(A.contiguous(), shuffle=True)
# Hardware-accelerated 4-bit GEMM using pre-shuffled B weights.
return aiter.gemm_a4w4(
A_q,
B_shuffle,
A_scale_sh,
B_scale_sh,
dtype=dtypes.bf16,
bpreshuffle=True,
)
scrolls · 40 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON