submission 609621
0xcouleurefil · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 58 lines, June 9 Researcher Reciprocity License v1.0.
submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-mxfp4-mm-609621?include=source"interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, mxfp4
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:9b38ae2203700a6a5005aad24fc18a830c5f50a131a0cec0a25764bb663982a3
license declaredunknown
license concludedunknown
authors0xcouleurefil
imported2026-08-26
Kernel source
submission.py58 lines
import torch
import aiter
from task import input_t, output_t
from aiter import dtypes
from aiter.ops.triton.quant import dynamic_mxfp4_quant
from aiter.utility.fp4_utils import e8m0_shuffle
def _quant_mxfp4(x, shuffle=True):
x_fp4, bs_e8m0 = dynamic_mxfp4_quant(x)
if shuffle:
bs_e8m0 = e8m0_shuffle(bs_e8m0)
return x_fp4.view(dtypes.fp4x2), bs_e8m0.view(dtypes.fp8_e8m0)
# 🛡️ 终极防御路由:大矩阵提速,小矩阵保命
def safe_router(M, N, K):
cfg = {
"BLOCK_SIZE_M": 128,
"BLOCK_SIZE_N": 256,
"BLOCK_SIZE_K": 64,
"num_warps": 8,
"num_stages": 3,
"NUM_KSPLIT": 1,
"waves_per_eu": 2,
"matrix_instr_nonkdim": 32
}
# 💥 核心修复:如果 M 太小,坚决不能开启 GROUP_SIZE_M,否则必炸!
if M <= 128:
cfg["GROUP_SIZE_M"] = 1
else:
cfg["GROUP_SIZE_M"] = 4
return cfg, None
def custom_kernel(data: input_t) -> output_t:
A, B, B_q, B_shuffle, B_scale_sh = data
A_contig = A.contiguous()
A_q, A_scale_sh = _quant_mxfp4(A_contig, shuffle=True)
# 绝对安全的内存对齐
B_shuffle_contig = B_shuffle.contiguous()
B_scale_sh_contig = B_scale_sh.contiguous()
# 注入智能防御路由
from aiter.ops.triton._triton_kernels.gemm.basic import gemm_afp4wfp4
gemm_afp4wfp4._get_config = safe_router
out_gemm = aiter.gemm_a4w4(
A_q,
B_shuffle_contig,
A_scale_sh,
B_scale_sh_contig,
dtype=dtypes.bf16,
bpreshuffle=True
)
return out_gemm
scrolls · 58 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON