submission 594954
Shlok · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 163 lines, June 9 Researcher Reciprocity License v1.0.
amd-moe-mxfp4.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-moe-mxfp4-594954?include=source"interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, fp32, fp8_e8m0, int32, mxfp4
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:e21580c623a8466bd8c3972e70170adc759ddbdc5dc7acd0a4b19560c07e678a
license declaredunknown
license concludedunknown
authorsShlok
imported2026-08-26
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
fp4
Submission template for DeepSeek-R1 MXFP4 MoE kernel.Kernel source
amd-moe-mxfp4.py163 lines
#!POPCORN leaderboard amd-moe-mxfp4
# This is a submission template for popcorn leaderboard 'amd-moe-mxfp4'.
# Your task is as follows:
# > You will implement a DeepSeek-R1 style MXFP4 Mixture-of-Experts (MoE) fused kernel optimized for AMD Instinct MI355X GPU.
# >
# > To be explicit, you will be given a tuple of tensors:
# > ```
# > (hidden_states,
# > gate_up_weight, down_weight, # fp4x2 raw
# > gate_up_weight_scale, down_weight_scale, # e8m0 raw
# > gate_up_weight_shuffled, down_weight_shuffled, # fp4x2 pre-shuffled
# > gate_up_weight_scale_shuffled, down_weight_scale_shuffled, # e8m0 pre-shuffled
# > topk_weights, topk_ids,
# > config)
# > ```
# > where:
# > * `hidden_states` is M x d_hidden in bfloat16 (the input activations, M = batch of tokens)
# > * `gate_up_weight` is [E, 2*d_expert_pad, d_hidden_pad//2] in MXFP4 (fp4x2), raw layout.
# > Fused gate + up projection weights for each expert. E = number of local experts.
# > * `down_weight` is [E, d_hidden_pad, d_expert_pad//2] in MXFP4 (fp4x2), raw layout.
# > Down projection weights for each expert.
# > * `gate_up_weight_scale` is [E, 2*d_expert_pad, d_hidden_pad//32] in E8M0, raw layout.
# > Block scales (block_size=32) for gate_up_weight.
# > * `down_weight_scale` is [E, d_hidden_pad, d_expert_pad//32] in E8M0, raw layout.
# > Block scales for down_weight.
# > * `gate_up_weight_shuffled` / `down_weight_shuffled` are the same weights shuffled to
# > (16,16) tile-coalesced layout for the CK kernel.
# > * `gate_up_weight_scale_shuffled` / `down_weight_scale_shuffled` are the scales after
# > e8m0_shuffle, flattened to [padded, flat].
# > * `topk_weights` is [M, total_top_k] float32: routing weights (routed experts + shared experts).
# > * `topk_ids` is [M, total_top_k] int32: expert indices. First nexpertspertoken columns are
# > routed expert ids (0..n_routed-1), last nsharedexperts columns are shared expert ids
# > (n_routed..n_routed+n_shared-1). Shared experts are always selected with weight=1.0.
# > * `config` is a dict with: d_hidden, d_expert, d_hidden_pad, d_expert_pad,
# > n_routed_experts, n_shared_experts, n_experts_per_token, total_top_k, bs.
# >
# > Then, the fused_moe kernel flow is:
# > (1) Quant activations to MXFP4: aiter per-1x32 dynamic quantization of hidden_states.
# > (2) Stage 1 GEMM + activation (per token i, per assigned expert j):
# > - gate = x_i @ W_gate_j.T # [d_hidden] x [d_expert, d_hidden].T -> [d_expert]
# > - up = x_i @ W_up_j.T # [d_hidden] x [d_expert, d_hidden].T -> [d_expert]
# > - intermediate = SiLU(gate) * up # SwiGLU activation, -> [d_expert]
# > (W_gate and W_up are fused as gate_up_weight, so this is one a4w4 GEMM + fused activation)
# > (3) Stage 2 GEMM:
# > - expert_out = intermediate @ W_down_j.T # [d_expert] x [d_hidden, d_expert].T -> [d_hidden]
# > (4) Weighted reduction:
# > - output_i += w_ij * expert_out # accumulate across top_k experts
# > All weight GEMMs are a4w4 (MXFP4 activations x MXFP4 weights, per-1x32 block scaling).
# > The AITER CK kernel fuses all of the above into a 2-stage pipeline across all tokens and experts.
# >
# > DeepSeek-R1 MoE specs:
# > - hidden_size = 7168, moe_intermediate_size = 2048
# > - 256 routed experts + 1 shared expert (total 257), top-8 routed + 1 shared = 9 per token
# > - 58 MoE layers (layer 3-60)
# > - The shared expert processes ALL tokens unconditionally (weight=1.0)
# >
# > d_hidden_pad and d_expert_pad are the dimensions padded to 256-alignment for the CK kernel.
# >
# > The ranking criteria is the geometric mean of the benchmark results.
# >
# > ```
# > The AITER reference performance is (E includes shared expert, top_k = routed + shared):
# > bs E d_hidden d_expert top_k time[us]
# > 16 257 7168 256 9 152.7
# > 128 257 7168 256 9 239.0
# > 512 257 7168 256 9 336.5
# > 16 33 7168 512 9 106.2
# > 128 33 7168 512 9 141.1
# > 512 33 7168 512 9 225.0
# > 512 33 7168 2048 9 380.4
# > ```
# >
# > Input:
# > - hidden_states: [M, d_hidden] bf16
# > - gate_up_weight: [E, 2*d_expert_pad, d_hidden_pad//2] fp4x2 (raw, before shuffle)
# > - down_weight: [E, d_hidden_pad, d_expert_pad//2] fp4x2 (raw, before shuffle)
# > - gate_up_weight_scale: [E, 2*d_expert_pad, d_hidden_pad//32] e8m0 (raw, before shuffle)
# > - down_weight_scale: [E, d_hidden_pad, d_expert_pad//32] e8m0 (raw, before shuffle)
# > - gate_up_weight_shuffled: [E, 2*d_expert_pad, d_hidden_pad//2] fp4x2 (pre-shuffled for CK)
# > - down_weight_shuffled: [E, d_hidden_pad, d_expert_pad//2] fp4x2 (pre-shuffled for CK)
# > - gate_up_weight_scale_shuffled: [padded, flat] e8m0 (pre-shuffled for CK)
# > - down_weight_scale_shuffled: [padded, flat] e8m0 (pre-shuffled for CK)
# > - topk_weights: [M, total_top_k] float32
# > - topk_ids: [M, total_top_k] int32
# > - config: dict with dimensions
# >
# > Output:
# > - output: [M, d_hidden] bf16
# The deadline for this leaderboard is 2026-04-07 07:59:00+00:00
# You can automatically route this file to specific GPUs by adding a line
# `#!POPCORN gpus <GPUs>` to the header of this file.
# Happy hacking!
import torch
from typing import Dict
from task import input_t, output_t
from aiter import ActivationType, QuantType
from aiter.fused_moe import fused_moe
def custom_kernel(data: input_t) -> output_t:
"""
Submission template for DeepSeek-R1 MXFP4 MoE kernel.
Input data tuple:
hidden_states: [M, d_hidden] bf16
gate_up_weight: [E, 2*d_expert_pad, d_hidden_pad//2] fp4x2 (raw)
down_weight: [E, d_hidden_pad, d_expert_pad//2] fp4x2 (raw)
gate_up_weight_scale: [E, 2*d_expert_pad, scale_K] e8m0 (raw)
down_weight_scale: [E, d_hidden_pad, scale_K] e8m0 (raw)
gate_up_weight_shuffled: [E, 2*d_expert_pad, d_hidden_pad//2] fp4x2 (shuffled)
down_weight_shuffled: [E, d_hidden_pad, d_expert_pad//2] fp4x2 (shuffled)
gate_up_weight_scale_shuffled:[padded, flat] e8m0 (shuffled)
down_weight_scale_shuffled: [padded, flat] e8m0 (shuffled)
topk_weights: [M, total_top_k] float32
topk_ids: [M, total_top_k] int32
config: dict
Returns:
output: [M, d_hidden] bf16
"""
(
hidden_states,
gate_up_weight,
down_weight,
gate_up_weight_scale,
down_weight_scale,
gate_up_weight_shuffled,
down_weight_shuffled,
gate_up_weight_scale_shuffled,
down_weight_scale_shuffled,
topk_weights,
topk_ids,
config,
) = data
hidden_pad = config["d_hidden_pad"] - config["d_hidden"]
intermediate_pad = config["d_expert_pad"] - config["d_expert"]
output = fused_moe(
hidden_states,
gate_up_weight_shuffled,
down_weight_shuffled,
topk_weights,
topk_ids,
expert_mask=None,
activation=ActivationType.Silu,
quant_type=QuantType.per_1x32,
doweight_stage1=False,
w1_scale=gate_up_weight_scale_shuffled,
w2_scale=down_weight_scale_shuffled,
a1_scale=None,
a2_scale=None,
hidden_pad=hidden_pad,
intermediate_pad=intermediate_pad,
)
return output
scrolls · 163 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON