submission 646755
inference_and_chill · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 75 lines, June 9 Researcher Reciprocity License v1.0.
submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-moe-mxfp4-646755?include=source"interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, fp32, fp8_e8m0, int32, mxfp4
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:ddc0edf13655171d958305a088b5b4c2b6b1de695142fa4ae4173827f3bc29ea
license declaredunknown
license concludedunknown
authorsinference_and_chill
imported2026-08-26
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
Kernel source
submission.py75 lines
#!POPCORN leaderboard amd-moe-mxfp4
#!POPCORN gpu MI355X
"""MoE MXFP4: CKTile split-K for small-batch E=33 + tuned block_m.
Key optimization: monkey-patch get_ksplit to return split-K=4 ONLY for small
batch E=33 cases (M <= 128). Combined with is_shuffled=True, this enables the
CKTile MXFP4 split-K path which is 26% faster for bs=16/E=33.
Large batch (M=512) uses default ksplit=0 (no split-K).
E=257 uses CSV-tuned CK kernels (unaffected by get_ksplit).
"""
from __future__ import annotations
import functools
# Monkey-patch get_ksplit BEFORE importing fused_moe
import aiter.fused_moe as _fm
@functools.lru_cache(maxsize=2048)
def _custom_get_ksplit(token, topk, expert, inter_dim, model_dim):
# CKTile split-K only helps small-batch E=33: 26% faster for bs=16, 5% for bs=128
# Hurts large batch (bs=512): 45-155% slower. So only enable for small M.
if token <= 128 and expert <= 64:
return 4
return 0
_fm.get_ksplit = _custom_get_ksplit
from aiter import ActivationType, QuantType
from aiter.fused_moe import fused_moe
from task import input_t, output_t
# block_m overrides for E=33 (CKTile supports {16, 32, 64})
_BLOCK_M = {
(33, 512, 16): 32,
(33, 512, 128): 32,
(33, 512, 512): 64,
(33, 2048, 512): 64,
}
def custom_kernel(data: input_t) -> output_t:
(
hidden_states, w1_raw, w2_raw, w1s_raw, w2s_raw,
w1_shuf, w2_shuf, w1ss_shuf, w2ss_shuf,
topk_weights, topk_ids, config,
) = data
hidden_pad = int(config["d_hidden_pad"]) - int(config["d_hidden"])
intermediate_pad = int(config["d_expert_pad"]) - int(config["d_expert"])
M = hidden_states.shape[0]
E = int(config["n_routed_experts"]) + int(config["n_shared_experts"])
d_expert_pad = int(config["d_expert_pad"])
block_m = _BLOCK_M.get((E, d_expert_pad, M))
# Enable CKTile split-K for small batch (matches get_ksplit condition above)
use_cktile = (E <= 64 and M <= 128)
w1_shuf.is_shuffled = use_cktile
w2_shuf.is_shuffled = use_cktile
return fused_moe(
hidden_states, w1_shuf, w2_shuf, topk_weights, topk_ids,
activation=ActivationType.Silu,
quant_type=QuantType.per_1x32,
w1_scale=w1ss_shuf,
w2_scale=w2ss_shuf,
hidden_pad=hidden_pad,
intermediate_pad=intermediate_pad,
block_size_M=block_m,
)
scrolls · 75 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON