submission 694654
DiegoCao · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 63 lines, June 9 Researcher Reciprocity License v1.0.
submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-moe-mxfp4-694654?include=source"interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, fp32, fp8_e8m0, int32, mxfp4
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:73fe1f7e737fefbfbab4829c943989112991757f64ff709c2f916e2764020b0f
license declaredunknown
license concludedunknown
authorsDiegoCao
imported2026-08-15
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
Kernel source
submission.py63 lines
#!POPCORN leaderboard amd-moe-mxfp4
#!POPCORN gpu MI355X
"""
MoE MXFP4 v13: Adaptive splitk via env var manipulation + cache clearing.
The fused_moe's get_ksplit uses @lru_cache keyed on (token, topk, expert, inter_dim, model_dim).
Since different workloads have different (token, expert, inter_dim), each gets its own
cache entry. The AITER_KSPLIT env var is read INSIDE get_ksplit on cache miss.
Strategy: set AITER_KSPLIT=2 for low-occupancy workloads, 0 for others.
Since each workload size is unique, the cache will miss on first call for each size.
We just need to set the env var BEFORE the first call to fused_moe for each workload.
"""
import os
import torch
from task import input_t, output_t
os.environ.setdefault("AITER_USE_OPUS_MOE_SORTING", "1")
from aiter import ActivationType, QuantType
from aiter.fused_moe import fused_moe
_SILU = ActivationType.Silu
_PER_1x32 = QuantType.per_1x32
_fused_moe = fused_moe
def custom_kernel(data: input_t) -> output_t:
(
hidden_states,
_, _, _, _,
w1, w2,
w1_scale, w2_scale,
topk_weights,
topk_ids,
config,
) = data
M = hidden_states.shape[0]
topk = topk_ids.shape[1]
E = w1.shape[0]
# Adaptive splitk: set env var before fused_moe's lru_cache lookup
# Each unique (M, E, d_expert) will be a cache miss on first call
tokens_per_expert = M * topk / E
if tokens_per_expert < 64 and M <= 128:
os.environ["AITER_KSPLIT"] = "2"
else:
os.environ["AITER_KSPLIT"] = "0"
return _fused_moe(
hidden_states, w1, w2,
topk_weights, topk_ids,
activation=_SILU,
quant_type=_PER_1x32,
w1_scale=w1_scale,
w2_scale=w2_scale,
hidden_pad=config["d_hidden_pad"] - config["d_hidden"],
intermediate_pad=config["d_expert_pad"] - config["d_expert"],
)
scrolls · 63 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON