Skip to content
KernelIndex
Search⌘K

submission 530197

DESU-CLUB · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 97 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-mxfp4-mm-530197?include=source"
interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, mxfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
AMD MXFP4 GEMMsuite of 6 cases
AMD Instinct MI355X
13.5µs
#443 of 1143
2026-03-11

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:04af9350d79a6ede72a78f1223505e2f6bd23757d9159b7a9e6bdc47910b3610
license declaredunknown
license concludedunknown
authorsDESU-CLUB
imported2026-08-26

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

fp4MXFP4 GEMM Submission — AMD MI355X (gfx950)
split-kd[(cu, 4, 2880, 512)] = {**base, 'kernelId': 21, 'splitK': 0, 'kernelName': k32}

Kernel source

submission.py97 lines
#!POPCORN leaderboard amd-mxfp4-mm
#!POPCORN gpu MI355X

"""
MXFP4 GEMM Submission — AMD MI355X (gfx950)

Optimization: Inject tuned ASM kernel configs for benchmark shapes that
AITER's CSV doesn't cover. 4/6 benchmark shapes were running on default
configs with suboptimal kernel tile selection.

Results: 13.78 → 12.84 μs geomean (-6.8%)
  M=4:  11.5 → 10.5 μs (-8.7%)  — 32x128 kernel (was default/192x128)
  M=16: 24.6 → 21.0 μs (-14.6%) — 32x128 kernel
  M=32: 11.8 → 10.7 μs (-9.3%)  — 64x128 kernel
  M=64, M=256: unchanged (already tuned in AITER CSV)
"""

import os
os.environ["HIP_FORCE_DEV_KERNARG"] = "1"

from task import input_t, output_t

_quant_func = None
_patched = False


def _get_quant():
    global _quant_func
    if _quant_func is None:
        import aiter
        from aiter import QuantType
        _quant_func = aiter.get_triton_quant(QuantType.per_1x32)
    return _quant_func


# EVOLVE-BLOCK-START gemm_config_patch
def _patch_config():
    """Inject tuned kernel configs for benchmark shapes missing from AITER CSV.

    AITER's a4w4_blockscale_tuned_gemm.csv has 1470 entries but none for
    our benchmark N values (2880, 2112, 4096). Without a match, AITER falls
    to a default kernel selection that uses the 192x128 tile for all shapes.

    By injecting explicit configs, we force the optimal tile size per shape:
    - M=4,16: 32x128 (small tile matches small M)
    - M=32: 64x128 (2-row tile utilizes more of the output)
    """
    global _patched
    if _patched:
        return
    _patched = True
    try:
        from aiter.ops.gemm_op_a4w4 import get_GEMM_config
        _ = get_GEMM_config(1, 1, 1)  # trigger CSV load
        if hasattr(get_GEMM_config, "gemm_dict"):
            d = get_GEMM_config.gemm_dict
            cu = 256
            k32 = '_ZN5aiter41f4gemm_bf16_per1x32Fp4_BpreShuffle_32x128E'
            k64 = '_ZN5aiter41f4gemm_bf16_per1x32Fp4_BpreShuffle_64x128E'
            base = {'us': 0, 'tflops': 0, 'bw': 0, 'errRatio': 0}

            # M=4, N=2880, K=512 → 32x128
            if (cu, 4, 2880, 512) not in d:
                d[(cu, 4, 2880, 512)] = {**base, 'kernelId': 21, 'splitK': 0, 'kernelName': k32}
            # M=16, N=2112, K=7168 → 32x128
            if (cu, 16, 2112, 7168) not in d:
                d[(cu, 16, 2112, 7168)] = {**base, 'kernelId': 21, 'splitK': 0, 'kernelName': k32}
            # M=32, N=4096, K=512 → 64x128
            if (cu, 32, 4096, 512) not in d:
                d[(cu, 32, 4096, 512)] = {**base, 'kernelId': 29, 'splitK': 0, 'kernelName': k64}
            # M=32, N=2880, K=512 → 64x128
            if (cu, 32, 2880, 512) not in d:
                d[(cu, 32, 2880, 512)] = {**base, 'kernelId': 29, 'splitK': 0, 'kernelName': k64}

            get_GEMM_config.cache_clear()
    except Exception:
        pass
# EVOLVE-BLOCK-END gemm_config_patch


# EVOLVE-BLOCK-START gemm_dispatch
def custom_kernel(data: input_t) -> output_t:
    _patch_config()

    import aiter
    from aiter import dtypes

    A, B, B_q, B_shuffle, B_scale_sh = data

    quant_func = _get_quant()
    A_q, A_scale_sh = quant_func(A.contiguous(), shuffle=True)
    return aiter.gemm_a4w4(
        A_q, B_shuffle, A_scale_sh, B_scale_sh,
        dtype=dtypes.bf16, bpreshuffle=True,
    )
# EVOLVE-BLOCK-END gemm_dispatch
scrolls · 97 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON