submission 514005
ohamnl. · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 66 lines, June 9 Researcher Reciprocity License v1.0.
solution.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-mxfp4-mm-514005?include=source"interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, mxfp4
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:7edef0b75e522cd0016215287054e63665ddb7f5d19d8a22351594db6e4f1ba2
license declaredunknown
license concludedunknown
authorsohamnl.
imported2026-08-26
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
fp4
MXFP4 GEMM: bf16 A x MXFP4 B -> bf16 CKernel source
solution.py66 lines
#!POPCORN leaderboard amd-mxfp4-mm
#!POPCORN gpu MI355X
import torch
from task import input_t, output_t
from utils import make_match_reference
import aiter
from aiter import QuantType, dtypes
from aiter.ops.shuffle import shuffle_weight
# the ref quantizes A fresh every call with get_triton_quant
# that's the only real cost here since B is already pre-shuffled
# our job: make A quant + gemm as fast as possible
# cache the quant function — get_triton_quant does jit compilation on first call
# calling it once at module load means zero compilation overhead at runtime
_quant_func = None
def _get_quant():
global _quant_func
if _quant_func is None:
_quant_func = aiter.get_triton_quant(QuantType.per_1x32)
return _quant_func
def custom_kernel(data: input_t) -> output_t:
"""
MXFP4 GEMM: bf16 A x MXFP4 B -> bf16 C
ref_kernel quantizes A with get_triton_quant every single call.
We cache the quant function to avoid recompilation overhead.
B is already shuffled in generate_input — use it directly.
The real kernel is aiter.gemm_a4w4 which maps to CDNA4 native
MXFP4 tensor core instructions. We can't replace that.
What we can do: make sure everything feeding into it is optimal.
"""
A, B, B_q, B_shuffle, B_scale_sh = data
# contiguous check — gemm_a4w4 requires contiguous input
# in the ref this is called unconditionally, we only do it if needed
if not A.is_contiguous():
A = A.contiguous()
# quantize A to MXFP4 with shuffled layout
# shuffle=True matches what gemm_a4w4 expects with bpreshuffle=True
A_q, A_scale_sh = _get_quant()(A, shuffle=True)
# B_shuffle and B_scale_sh are pre-computed in generate_input
# zero extra work here vs ref which also uses them directly
out = aiter.gemm_a4w4(
A_q,
B_shuffle,
A_scale_sh,
B_scale_sh,
dtype=dtypes.bf16,
bpreshuffle=True,
)
return out
solution = custom_kernel
from reference import ref_kernel
check_implementation = make_match_reference(ref_kernel, rtol=1e-02, atol=1e-02)scrolls · 66 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON