submission 583919
Zobin Huang · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 70 lines, June 9 Researcher Reciprocity License v1.0.
submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-moe-mxfp4-583919?include=source"interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, fp32, fp8_e8m0, int32, mxfp4
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:8d124920c2a49cde0f5c8019c3d56f357c1871d4178d0b37d29bea6956d0c7ab
license declaredunknown
license concludedunknown
authorsZobin Huang
imported2026-08-26
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
fp4
"""Optimized MoE MXFP4 submission for AMD Instinct MI355X.Kernel source
submission.py70 lines
"""Optimized MoE MXFP4 submission for AMD Instinct MI355X.
Uses AITER's fused_moe with per-config block_size_M tuning.
The block_size_M controls the tile size for the CK/ASM kernel dispatch:
- For many-expert configs (E=257): smaller block_m=32 reduces padding waste
- For fewer-expert configs (E=33): larger block_m balances CU utilization
"""
from task import input_t, output_t
import torch
from aiter import ActivationType, QuantType
from aiter.fused_moe import fused_moe
def custom_kernel(data: input_t) -> output_t:
(
hidden_states,
gate_up_weight,
down_weight,
gate_up_weight_scale,
down_weight_scale,
gate_up_weight_shuffled,
down_weight_shuffled,
gate_up_weight_scale_shuffled,
down_weight_scale_shuffled,
topk_weights,
topk_ids,
config,
) = data
hidden_pad = config["d_hidden_pad"] - config["d_hidden"]
intermediate_pad = config["d_expert_pad"] - config["d_expert"]
# Select block_size_M based on problem shape
E_total = config["n_routed_experts"] + config["n_shared_experts"]
bs = config["bs"]
d_expert = config["d_expert"]
# For many-expert configs, block_m=32 minimizes padding waste
# For fewer-expert configs, heuristic auto-selection is usually good
if E_total > 64:
block_m = 32
elif d_expert >= 2048:
block_m = 64
elif bs >= 128:
block_m = 128
else:
block_m = 64
output = fused_moe(
hidden_states,
gate_up_weight_shuffled,
down_weight_shuffled,
topk_weights,
topk_ids,
expert_mask=None,
activation=ActivationType.Silu,
quant_type=QuantType.per_1x32,
doweight_stage1=False,
w1_scale=gate_up_weight_scale_shuffled,
w2_scale=down_weight_scale_shuffled,
a1_scale=None,
a2_scale=None,
block_size_M=block_m,
hidden_pad=hidden_pad,
intermediate_pad=intermediate_pad,
)
return output
scrolls · 70 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON