submission 530197
DESU-CLUB · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 97 lines, June 9 Researcher Reciprocity License v1.0.
submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-mxfp4-mm-530197?include=source"interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, mxfp4
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:04af9350d79a6ede72a78f1223505e2f6bd23757d9159b7a9e6bdc47910b3610
license declaredunknown
license concludedunknown
authorsDESU-CLUB
imported2026-08-26
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
Kernel source
submission.py97 lines
#!POPCORN leaderboard amd-mxfp4-mm
#!POPCORN gpu MI355X
"""
MXFP4 GEMM Submission — AMD MI355X (gfx950)
Optimization: Inject tuned ASM kernel configs for benchmark shapes that
AITER's CSV doesn't cover. 4/6 benchmark shapes were running on default
configs with suboptimal kernel tile selection.
Results: 13.78 → 12.84 μs geomean (-6.8%)
M=4: 11.5 → 10.5 μs (-8.7%) — 32x128 kernel (was default/192x128)
M=16: 24.6 → 21.0 μs (-14.6%) — 32x128 kernel
M=32: 11.8 → 10.7 μs (-9.3%) — 64x128 kernel
M=64, M=256: unchanged (already tuned in AITER CSV)
"""
import os
os.environ["HIP_FORCE_DEV_KERNARG"] = "1"
from task import input_t, output_t
_quant_func = None
_patched = False
def _get_quant():
global _quant_func
if _quant_func is None:
import aiter
from aiter import QuantType
_quant_func = aiter.get_triton_quant(QuantType.per_1x32)
return _quant_func
# EVOLVE-BLOCK-START gemm_config_patch
def _patch_config():
"""Inject tuned kernel configs for benchmark shapes missing from AITER CSV.
AITER's a4w4_blockscale_tuned_gemm.csv has 1470 entries but none for
our benchmark N values (2880, 2112, 4096). Without a match, AITER falls
to a default kernel selection that uses the 192x128 tile for all shapes.
By injecting explicit configs, we force the optimal tile size per shape:
- M=4,16: 32x128 (small tile matches small M)
- M=32: 64x128 (2-row tile utilizes more of the output)
"""
global _patched
if _patched:
return
_patched = True
try:
from aiter.ops.gemm_op_a4w4 import get_GEMM_config
_ = get_GEMM_config(1, 1, 1) # trigger CSV load
if hasattr(get_GEMM_config, "gemm_dict"):
d = get_GEMM_config.gemm_dict
cu = 256
k32 = '_ZN5aiter41f4gemm_bf16_per1x32Fp4_BpreShuffle_32x128E'
k64 = '_ZN5aiter41f4gemm_bf16_per1x32Fp4_BpreShuffle_64x128E'
base = {'us': 0, 'tflops': 0, 'bw': 0, 'errRatio': 0}
# M=4, N=2880, K=512 → 32x128
if (cu, 4, 2880, 512) not in d:
d[(cu, 4, 2880, 512)] = {**base, 'kernelId': 21, 'splitK': 0, 'kernelName': k32}
# M=16, N=2112, K=7168 → 32x128
if (cu, 16, 2112, 7168) not in d:
d[(cu, 16, 2112, 7168)] = {**base, 'kernelId': 21, 'splitK': 0, 'kernelName': k32}
# M=32, N=4096, K=512 → 64x128
if (cu, 32, 4096, 512) not in d:
d[(cu, 32, 4096, 512)] = {**base, 'kernelId': 29, 'splitK': 0, 'kernelName': k64}
# M=32, N=2880, K=512 → 64x128
if (cu, 32, 2880, 512) not in d:
d[(cu, 32, 2880, 512)] = {**base, 'kernelId': 29, 'splitK': 0, 'kernelName': k64}
get_GEMM_config.cache_clear()
except Exception:
pass
# EVOLVE-BLOCK-END gemm_config_patch
# EVOLVE-BLOCK-START gemm_dispatch
def custom_kernel(data: input_t) -> output_t:
_patch_config()
import aiter
from aiter import dtypes
A, B, B_q, B_shuffle, B_scale_sh = data
quant_func = _get_quant()
A_q, A_scale_sh = quant_func(A.contiguous(), shuffle=True)
return aiter.gemm_a4w4(
A_q, B_shuffle, A_scale_sh, B_scale_sh,
dtype=dtypes.bf16, bpreshuffle=True,
)
# EVOLVE-BLOCK-END gemm_dispatch
scrolls · 97 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON