Skip to content
KernelIndex
Search⌘K

submission 681091

hwk603 · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 40 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-mxfp4-mm-681091?include=source"
interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, mxfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
AMD MXFP4 GEMMsuite of 6 cases
AMD Instinct MI355X
22.6µs
#762 of 1143
2026-03-31

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:04732fff3e6d87d49d167bf59e204f57ce04dab4408526fe07e98734b089a17c
license declaredunknown
license concludedunknown
authorshwk603
imported2026-08-26

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

fp4Optimized MXFP4 GEMM: module-level imports + config patch for missing shapes.
split-kd[key] = {'kernelId':21,'splitK':0,'us':0,'kernelName':_K32,'tflops':0,'bw':0,'errRatio':0.0}

Kernel source

submission.py40 lines
"""
Optimized MXFP4 GEMM: module-level imports + config patch for missing shapes.
"""
import torch
import aiter
from aiter import dtypes
from aiter.ops.triton.quant import dynamic_mxfp4_quant
from aiter.utility.fp4_utils import e8m0_shuffle
from aiter.ops.gemm_op_a4w4 import get_GEMM_config
from aiter.jit.utils.chip_info import get_cu_num

from task import input_t, output_t

_K32 = "_ZN5aiter41f4gemm_bf16_per1x32Fp4_BpreShuffle_32x128E"

def _patch_configs():
    get_GEMM_config(1, 512, 4096)
    d = get_GEMM_config.gemm_dict
    cu = get_cu_num()
    for key in [(cu,4,2880,512),(cu,16,2112,7168),(cu,32,4096,512),(cu,32,2880,512)]:
        if key not in d:
            d[key] = {'kernelId':21,'splitK':0,'us':0,'kernelName':_K32,'tflops':0,'bw':0,'errRatio':0.0}

_patch_configs()


def custom_kernel(data: input_t) -> output_t:
    A, B, B_q, B_shuffle, B_scale_sh = data
    m, k = A.shape

    A_fp4, A_scale = dynamic_mxfp4_quant(A)
    A_scale_sh = e8m0_shuffle(A_scale)
    A_q = A_fp4.view(dtypes.fp4x2)
    A_scale_sh = A_scale_sh.view(dtypes.fp8_e8m0)

    return aiter.gemm_a4w4(
        A_q, B_shuffle, A_scale_sh, B_scale_sh,
        dtype=dtypes.bf16, bpreshuffle=True,
    )
scrolls · 40 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON