submission 594929
Shlok · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 81 lines, June 9 Researcher Reciprocity License v1.0.
amd-mxfp4-mm.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-mxfp4-mm-594929?include=source"interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, mxfp4
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:cb627ad330a0972d049e707678e1e420e6ac6227699e954abd45005152f228ce
license declaredunknown
license concludedunknown
authorsShlok
imported2026-08-26
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
fp4
FP4 quant + FP4 GEMM reference: bf16 A, MXFP4 B -> MXFP4 per-1x32 quant A -> gemm_a4w4 -> bf16 C.Kernel source
amd-mxfp4-mm.py81 lines
#!POPCORN leaderboard amd-mxfp4-mm
# This is a submission template for popcorn leaderboard 'amd-mxfp4-mm'.
# Your task is as follows:
# > You will implement a quantize func and block scaled MXFP4 matrix-matrix multiplication kernel optimized for AMD Instinct MI355X GPU.
# > To be explicit, you will be given a tuple of tensors:
# > ```
# > (A, B, B_q, B_shuffle, B_scale_sh)
# > ```
# > where:
# > * `A` is M x K in K-major order in bfloat16
# > * `B` is N x K in K-major order in bfloat16
# > * `B_q` is N x K/2 in K-major order in MXFP4
# > * `B_shuffle` is N x K/2 in shuffled order in MXFP4, shuffled to (16,16) tile coalesced
# > * `B_scale_sh` is * x K/32 in E8M0, * means it will be padded.
# >
# > Then, the kernel flow is bf16 A, MXFP4 B -> MXFP4 per-1x32 quant A -> gemm_a4w4 -> BF16 C [m,n].
# > To be specific, the invocation flow is:
# > (1) Quant A to MXFP4: aiter.get_triton_quant(QuantType.per_1x32).
# > (2) GEMM: aiter.gemm_a4w4.
# > m, n divisible by 64; k divisible by 64.
# >
# > The ranking criteria is the geometric mean of the benchmark results.
# > Pls note that this is the elimination round, whoever rank top5 are selected into the next round, e2e optimization for deepseek-R1-MXFP4 and GPTOSS-MXFP4 mdoel
# > ```
# > The aiter performance is:
# > M N K time[us]
# > 4 2880 512 8.198
# > 16 2112 7168 20.873
# > 32 4096 512 9.462
# > 32 2880 512 9.173
# > 64 7168 2048 12.738
# > 256 3072 1536 12.219
# > ```
# The deadline for this leaderboard is 2026-04-07 07:59:00+00:00
# You can automatically route this file to specific GPUs by adding a line
# `#!POPCORN gpus <GPUs>` to the header of this file.
# Happy hacking!
"""
FP4 quant + FP4 GEMM reference: bf16 A, MXFP4 B -> MXFP4 per-1x32 quant A -> gemm_a4w4 -> bf16 C.
Quant logic follows aiter op_tests/test_gemm_a4w4.py (get_triton_quant(QuantType.per_1x32)).
"""
from task import input_t, output_t
def custom_kernel(data: input_t) -> output_t:
"""
Reference: MXFP4 per-1x32 quant on A; B_shuffle, B_scale_sh from generate_input.
gemm_a4w4 with bpreshuffle=True.
"""
import aiter
from aiter import QuantType, dtypes
from aiter.ops.triton.quant import dynamic_mxfp4_quant
from aiter.utility.fp4_utils import e8m0_shuffle
def _quant_mxfp4(x, shuffle=True):
x_fp4, bs_e8m0 = dynamic_mxfp4_quant(x)
if shuffle:
bs_e8m0 = e8m0_shuffle(bs_e8m0)
return x_fp4.view(dtypes.fp4x2), bs_e8m0.view(dtypes.fp8_e8m0)
A, B, B_q, B_shuffle, B_scale_sh = data
A = A.contiguous()
B = B.contiguous()
m, k = A.shape
n, _ = B.shape
A_q, A_scale_sh = _quant_mxfp4(A, shuffle=True)
out_gemm = aiter.gemm_a4w4(
A_q,
B_shuffle,
A_scale_sh,
B_scale_sh,
dtype=dtypes.bf16,
bpreshuffle=True,
)
return out_gemm
scrolls · 81 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON