submission 718275
jin051623 · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 66 lines, June 9 Researcher Reciprocity License v1.0.
submission1.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-mxfp4-mm-718275?include=source"interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, mxfp4
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:256d1d12bc84f795b11aa22ce06bd3ce5b5c3498c125d8d9e1057fb0afd70a44
license declaredunknown
license concludedunknown
authorsjin051623
imported2026-08-26
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
fp4
FP4 quant + FP4 GEMM reference: bf16 A, MXFP4 B -> MXFP4 per-1x32 quant A -> gemm_a4w4 -> bf16 C.Kernel source
submission1.py66 lines
#!POPCORN leaderboard amd-mxfp4-mm
#!POPCORN gpu MI355X
"""
Optimized submission: same semantics as reference but lower Python overhead.
FP4 quant + FP4 GEMM reference: bf16 A, MXFP4 B -> MXFP4 per-1x32 quant A -> gemm_a4w4 -> bf16 C.
"""
from task import input_t, output_t
# Module-scope imports to avoid per-call overhead
import aiter
from aiter import dtypes
from aiter.ops.triton.quant import dynamic_mxfp4_quant # patched kernel per reference
from aiter.utility.fp4_utils import e8m0_shuffle
# Keep SCALE_GROUP_SIZE visible if needed by future local checks
SCALE_GROUP_SIZE = 32
# Module-scope helper: quantizes bf16 -> packed fp4x2 + shuffled E8M0 scales
def _quant_mxfp4_per_1x32(x, shuffle: bool = True):
"""
x: bf16 [M, K]
returns:
- x_fp4_packed: dtypes.fp4x2 view (packed)
- bs_e8m0_shuffled: dtypes.fp8_e8m0 view (shuffled if requested)
"""
# dynamic_mxfp4_quant is the patched, Triton-backed quantizer referenced in the tests
x_fp4, bs_e8m0 = dynamic_mxfp4_quant(x)
if shuffle:
bs_e8m0 = e8m0_shuffle(bs_e8m0)
# Return views in the exact dtype shapes expected by gemm_a4w4
return x_fp4.view(dtypes.fp4x2), bs_e8m0.view(dtypes.fp8_e8m0)
def custom_kernel(data: input_t) -> output_t:
"""
Optimized hot-path:
- Quantize A (bf16) to MXFP4 per-1x32 (packed)
- Call aiter.gemm_a4w4 with pre-shuffled B and scales
- Return bf16 C
"""
# Unpack only what we need; keep variable names consistent with reference harness
A, B, B_q, B_shuffle, B_scale_sh = data
# Ensure A is contiguous once (avoid unconditional copy if already contiguous)
if not A.is_contiguous():
A = A.contiguous()
# Quantize A to MXFP4 per-1x32 and get shuffled scales
A_q, A_scale_sh = _quant_mxfp4_per_1x32(A, shuffle=True)
# Call the Triton-backed GEMM (keeps correctness and uses optimized backend)
# bpreshuffle=True matches the reference harness expectations
out = aiter.gemm_a4w4(
A_q,
B_shuffle,
A_scale_sh,
B_scale_sh,
dtype=dtypes.bf16,
bpreshuffle=True,
)
return out
scrolls · 66 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON