Skip to content
KernelIndex
Search⌘K

submission 646755

inference_and_chill · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 75 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-moe-mxfp4-646755?include=source"
interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, fp32, fp8_e8m0, int32, mxfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
AMD MXFP4 MoEsuite of 7 cases
AMD Instinct MI355X
182.5µs
#518 of 782
2026-03-27

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:ddc0edf13655171d958305a088b5b4c2b6b1de695142fa4ae4173827f3bc29ea
license declaredunknown
license concludedunknown
authorsinference_and_chill
imported2026-08-26

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

fp4"""MoE MXFP4: CKTile split-K for small-batch E=33 + tuned block_m.
split-k"""MoE MXFP4: CKTile split-K for small-batch E=33 + tuned block_m.

Kernel source

submission.py75 lines
#!POPCORN leaderboard amd-moe-mxfp4
#!POPCORN gpu MI355X

"""MoE MXFP4: CKTile split-K for small-batch E=33 + tuned block_m.

Key optimization: monkey-patch get_ksplit to return split-K=4 ONLY for small
batch E=33 cases (M <= 128). Combined with is_shuffled=True, this enables the
CKTile MXFP4 split-K path which is 26% faster for bs=16/E=33.

Large batch (M=512) uses default ksplit=0 (no split-K).
E=257 uses CSV-tuned CK kernels (unaffected by get_ksplit).
"""

from __future__ import annotations

import functools

# Monkey-patch get_ksplit BEFORE importing fused_moe
import aiter.fused_moe as _fm

@functools.lru_cache(maxsize=2048)
def _custom_get_ksplit(token, topk, expert, inter_dim, model_dim):
    # CKTile split-K only helps small-batch E=33: 26% faster for bs=16, 5% for bs=128
    # Hurts large batch (bs=512): 45-155% slower. So only enable for small M.
    if token <= 128 and expert <= 64:
        return 4
    return 0

_fm.get_ksplit = _custom_get_ksplit

from aiter import ActivationType, QuantType
from aiter.fused_moe import fused_moe
from task import input_t, output_t

# block_m overrides for E=33 (CKTile supports {16, 32, 64})
_BLOCK_M = {
    (33, 512, 16): 32,
    (33, 512, 128): 32,
    (33, 512, 512): 64,
    (33, 2048, 512): 64,
}


def custom_kernel(data: input_t) -> output_t:
    (
        hidden_states, w1_raw, w2_raw, w1s_raw, w2s_raw,
        w1_shuf, w2_shuf, w1ss_shuf, w2ss_shuf,
        topk_weights, topk_ids, config,
    ) = data

    hidden_pad = int(config["d_hidden_pad"]) - int(config["d_hidden"])
    intermediate_pad = int(config["d_expert_pad"]) - int(config["d_expert"])

    M = hidden_states.shape[0]
    E = int(config["n_routed_experts"]) + int(config["n_shared_experts"])
    d_expert_pad = int(config["d_expert_pad"])

    block_m = _BLOCK_M.get((E, d_expert_pad, M))

    # Enable CKTile split-K for small batch (matches get_ksplit condition above)
    use_cktile = (E <= 64 and M <= 128)
    w1_shuf.is_shuffled = use_cktile
    w2_shuf.is_shuffled = use_cktile

    return fused_moe(
        hidden_states, w1_shuf, w2_shuf, topk_weights, topk_ids,
        activation=ActivationType.Silu,
        quant_type=QuantType.per_1x32,
        w1_scale=w1ss_shuf,
        w2_scale=w2ss_shuf,
        hidden_pad=hidden_pad,
        intermediate_pad=intermediate_pad,
        block_size_M=block_m,
    )
scrolls · 75 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON