submission 697462
Renjie-gif · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 67 lines, June 9 Researcher Reciprocity License v1.0.
submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-mxfp4-mm-697462?include=source"interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, mxfp4
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:6c04cde41ec088c91bb2bab70d22497aec8704d48e00fe1f94c3977d943501cf
license declaredunknown
license concludedunknown
authorsRenjie-gif
imported2026-08-26
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
Kernel source
submission.py67 lines
"""
MXFP4 GEMM — Best: direct asm + quant caching + per-shape splitK + optimal tile from CSV.
Now using tuned CSV data: for our benchmark shapes, the tuned config uses
32x128 (kernelId=21) for small M and 64x128 (kernelId=29) for some shapes.
"""
from task import input_t, output_t
import torch
from aiter import dtypes
from aiter.ops.triton.quant import dynamic_mxfp4_quant
from aiter.utility.fp4_utils import e8m0_shuffle
from aiter.ops.gemm_op_a4w4 import gemm_a4w4_asm
_fp4x2 = dtypes.fp4x2
_e8m0 = dtypes.fp8_e8m0
_last_did = None
_last_aq = None
_last_asc = None
# From tuned CSV: kernelId 21 = 32x128, kernelId 29 = 64x128
# For M<=32: 32x128 is standard
# For M=64: could use 64x128 (better M-dim utilization)
# For M=256: 192x128 (default) or try 256x128
_K32x128 = "_ZN5aiter41f4gemm_bf16_per1x32Fp4_BpreShuffle_32x128E"
_K64x128 = "_ZN5aiter41f4gemm_bf16_per1x32Fp4_BpreShuffle_64x128E"
_K128x128 = "_ZN5aiter42f4gemm_bf16_per1x32Fp4_BpreShuffle_128x128E"
_K192x128 = "_ZN5aiter42f4gemm_bf16_per1x32Fp4_BpreShuffle_192x128E"
_K256x128 = "_ZN5aiter42f4gemm_bf16_per1x32Fp4_BpreShuffle_256x128E"
def _get_config(m, n, k):
"""Per-shape optimal kernel + splitK from tuned CSV analysis."""
mp = (m + 31) // 32 * 32
if mp <= 32:
return _K32x128, 0 if k <= 512 else (2 if k <= 1536 else 3)
elif mp <= 64:
# 64x128 matches M=64 exactly — no wasted rows
return _K64x128, 0 if k <= 512 else (2 if k <= 1536 else 3)
elif mp <= 128:
return _K128x128, 0 if k <= 512 else (2 if k <= 1536 else 3)
elif mp <= 256:
# 192x128 gives 2 tiles for M=256 — better parallelism than 256x128 (1 tile)
return _K192x128, 0 if k <= 512 else (2 if k <= 1536 else 3)
else:
return _K256x128, 3
def custom_kernel(data: input_t) -> output_t:
global _last_did, _last_aq, _last_asc
A, B, _, B_shuffle, B_scale_sh = data
m, k = A.shape
n = B.shape[0]
did = (id(data), m, k)
if did != _last_did:
fp4, bs = dynamic_mxfp4_quant(A)
_last_aq = fp4.view(_fp4x2)
_last_asc = e8m0_shuffle(bs.view(_e8m0))
_last_did = did
mp = (m + 31) // 32 * 32
out = torch.empty((mp, n), dtype=torch.bfloat16, device=A.device)
kernel, sk = _get_config(m, n, k)
gemm_a4w4_asm(_last_aq.view(m, k // 2), B_shuffle, _last_asc, B_scale_sh,
out, kernel, None, 1.0, 0.0, True, sk)
return out[:m]
scrolls · 67 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON