Skip to content
KernelIndex
Search⌘K

submission 721128

yanchaomei · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 108 lines, June 9 Researcher Reciprocity License v1.0.

submission_moe.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-moe-mxfp4-721128?include=source"
interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, fp32, fp8_e8m0, int32, mxfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
AMD MXFP4 MoEsuite of 7 cases
AMD Instinct MI355X
160.6µs
#273 of 782
2026-04-04

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:5c15a1d659c3fcc3f207b1466dc2a579c5b5583973f42e600e8a52ddee0f05d2
license declaredunknown
license concludedunknown
authorsyanchaomei
imported2026-08-15

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

split-kactivation=activation, split_k=ksplit, dtype=dtype,

Kernel source

submission_moe.py108 lines
#!POPCORN leaderboard amd-moe-mxfp4
#!POPCORN gpu MI355X

"""
V44: Lean hybrid — maximum coverage CK Tile for small+medium batch,
CK 2-stage for large batch. NO pre-trigger (avoids lock contention).

CK Tile shapes (known faster from V19b/V22/V23/V24 data):
  E=257 bs=16:  sk=2, bm=32 → ~97µs (was 137µs with CK 2-stage)
  E=257 bs=128: sk=2, bm=32 → ~184µs (was 218µs)
  E=33  bs=16:  sk=4, bm=32 → ~67µs (was 95µs)
  E=33  bs=128: sk=2, bm=32 → ~129µs (marginal vs 130µs CK 2-stage)

CK 2-stage shapes (CK Tile was SLOWER in V19b):
  E=257 bs=512: CK 2-stage → ~249µs (CK Tile was ~350µs+)
  E=33  bs=512: CK 2-stage → ~209µs (CK Tile was ~350µs+)
  E=33  d=2048: CK 2-stage → ~349µs (CK Tile untested but likely worse)
"""

import torch
import os
import sys
import functools
from task import input_t, output_t

os.environ["AITER_USE_OPUS_MOE_SORTING"] = "1"

import aiter
from aiter import ActivationType, QuantType
import aiter.fused_moe as fmoe_mod
from aiter.fused_moe import (
    fused_moe, MOEMetadata,
    cktile_moe_stage1, cktile_moe_stage2,
)

print("[V44] Lean hybrid: CK Tile for small/med, CK 2-stage for large", file=sys.stderr)

_orig_cfgs = fmoe_mod.get_2stage_cfgs

@functools.lru_cache(maxsize=256)
def _v44_dispatch(token, model_dim, inter_dim, expert, topk,
                  dtype, q_dtype_a, q_dtype_w, q_type, use_g1u1,
                  activation, doweight_stage1, hidden_pad, intermediate_pad,
                  is_shuffled=True):
    orig = _orig_cfgs(token, model_dim, inter_dim, expert, topk,
                      dtype, q_dtype_a, q_dtype_w, q_type, use_g1u1,
                      activation, doweight_stage1, hidden_pad, intermediate_pad,
                      is_shuffled)

    if not (q_type == QuantType.per_1x32
            and activation == ActivationType.Silu
            and not doweight_stage1
            and is_shuffled):
        return orig

    use_cktile = False
    ksplit = 2
    block_m = 32

    # E=257: CK Tile for bs<=128
    if expert >= 128 and token <= 128:
        use_cktile = True
        ksplit = 2
    # E=33: CK Tile for bs<=16 (sk=4) and bs<=128 with d<=1024 (sk=2)
    elif expert < 128 and token <= 16:
        use_cktile = True
        ksplit = 4
    elif expert < 128 and token <= 128 and inter_dim <= 1024:
        use_cktile = True
        ksplit = 2
    # Everything else: CK 2-stage (NEVER CK Tile for bs=512)

    if use_cktile:
        print(f"[V44] CKTile: token={token} E={expert} inter={inter_dim} sk={ksplit} bm={block_m}",
              file=sys.stderr)
        return MOEMetadata(
            stage1=functools.partial(
                cktile_moe_stage1, n_pad_zeros=hidden_pad, k_pad_zeros=0,
                activation=activation, split_k=ksplit, dtype=dtype,
            ),
            stage2=functools.partial(
                cktile_moe_stage2, activation=activation,
                n_pad_zeros=intermediate_pad, k_pad_zeros=0,
            ),
            block_m=block_m, ksplit=ksplit, run_1stage=False,
            has_bias=False, use_non_temporal_load=orig.use_non_temporal_load,
        )

    return orig

fmoe_mod.get_2stage_cfgs = _v44_dispatch
print("[V44] Dispatch installed", file=sys.stderr)


def custom_kernel(data: input_t) -> output_t:
    (hidden_states, guw, dw, guws, dws, guw_sh, dw_sh,
     guws_sc_sh, dws_sc_sh, topk_weights, topk_ids, config) = data
    hp = config["d_hidden_pad"] - config["d_hidden"]
    ip = config["d_expert_pad"] - config["d_expert"]
    return fused_moe(
        hidden_states, guw_sh, dw_sh, topk_weights, topk_ids,
        expert_mask=None, activation=ActivationType.Silu,
        quant_type=QuantType.per_1x32, doweight_stage1=False,
        w1_scale=guws_sc_sh, w2_scale=dws_sc_sh,
        a1_scale=None, a2_scale=None,
        hidden_pad=hp, intermediate_pad=ip,
    )
scrolls · 108 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON