submission 518127
Danishlynx · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 31 lines, June 9 Researcher Reciprocity License v1.0.
submission_gemm_v01.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-mxfp4-mm-518127?include=source"interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, mxfp4
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:e608949bb66e9da42bb31f6909ad44081097e1b3923ebd885c5bdead101f0480
license declaredunknown
license concludedunknown
authorsDanishlynx
imported2026-08-15
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
fp4
MXFP4 GEMM submission v1 - AITER a4w4 baseline.Kernel source
submission_gemm_v01.py31 lines
"""
MXFP4 GEMM submission v1 - AITER a4w4 baseline.
Flow: bf16 A -> MXFP4 quant A -> gemm_a4w4(A_q, B_shuffled) -> bf16 C
"""
from task import input_t, output_t
def custom_kernel(data: input_t) -> output_t:
import aiter
from aiter import QuantType, dtypes
A, B, B_q, B_shuffle, B_scale_sh = data
A = A.contiguous()
m, k = A.shape
n = B.shape[0]
# Quantize A to MXFP4 with shuffle for CK kernel
quant_func = aiter.get_triton_quant(QuantType.per_1x32)
A_q, A_scale_sh = quant_func(A, shuffle=True)
# GEMM via AITER CK kernel (pre-shuffled weights)
out_gemm = aiter.gemm_a4w4(
A_q,
B_shuffle,
A_scale_sh,
B_scale_sh,
dtype=dtypes.bf16,
bpreshuffle=True,
)
return out_gemm
scrolls · 31 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON