Skip to content
KernelIndex
Search⌘K

submission 594954

Shlok · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 163 lines, June 9 Researcher Reciprocity License v1.0.

amd-moe-mxfp4.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-moe-mxfp4-594954?include=source"
interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, fp32, fp8_e8m0, int32, mxfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
AMD MXFP4 MoEsuite of 7 cases
AMD Instinct MI355X
179.9µs
#465 of 782
2026-03-20

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:e21580c623a8466bd8c3972e70170adc759ddbdc5dc7acd0a4b19560c07e678a
license declaredunknown
license concludedunknown
authorsShlok
imported2026-08-26

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

fp4Submission template for DeepSeek-R1 MXFP4 MoE kernel.

Kernel source

amd-moe-mxfp4.py163 lines
#!POPCORN leaderboard amd-moe-mxfp4

# This is a submission template for popcorn leaderboard 'amd-moe-mxfp4'.
# Your task is as follows:
# > You will implement a DeepSeek-R1 style MXFP4 Mixture-of-Experts (MoE) fused kernel optimized for AMD Instinct MI355X GPU.
# > 
# > To be explicit, you will be given a tuple of tensors:
# > ```
# > (hidden_states,
# >  gate_up_weight, down_weight,                                         # fp4x2 raw
# >  gate_up_weight_scale, down_weight_scale,                             # e8m0  raw
# >  gate_up_weight_shuffled, down_weight_shuffled,                       # fp4x2 pre-shuffled
# >  gate_up_weight_scale_shuffled, down_weight_scale_shuffled,           # e8m0  pre-shuffled
# >  topk_weights, topk_ids,
# >  config)
# > ```
# > where:
# > * `hidden_states` is M x d_hidden in bfloat16 (the input activations, M = batch of tokens)
# > * `gate_up_weight` is [E, 2*d_expert_pad, d_hidden_pad//2] in MXFP4 (fp4x2), raw layout.
# >   Fused gate + up projection weights for each expert. E = number of local experts.
# > * `down_weight` is [E, d_hidden_pad, d_expert_pad//2] in MXFP4 (fp4x2), raw layout.
# >   Down projection weights for each expert.
# > * `gate_up_weight_scale` is [E, 2*d_expert_pad, d_hidden_pad//32] in E8M0, raw layout.
# >   Block scales (block_size=32) for gate_up_weight.
# > * `down_weight_scale` is [E, d_hidden_pad, d_expert_pad//32] in E8M0, raw layout.
# >   Block scales for down_weight.
# > * `gate_up_weight_shuffled` / `down_weight_shuffled` are the same weights shuffled to
# >   (16,16) tile-coalesced layout for the CK kernel.
# > * `gate_up_weight_scale_shuffled` / `down_weight_scale_shuffled` are the scales after
# >   e8m0_shuffle, flattened to [padded, flat].
# > * `topk_weights` is [M, total_top_k] float32: routing weights (routed experts + shared experts).
# > * `topk_ids` is [M, total_top_k] int32: expert indices. First nexpertspertoken columns are
# >   routed expert ids (0..n_routed-1), last nsharedexperts columns are shared expert ids
# >   (n_routed..n_routed+n_shared-1). Shared experts are always selected with weight=1.0.
# > * `config` is a dict with: d_hidden, d_expert, d_hidden_pad, d_expert_pad,
# >   n_routed_experts, n_shared_experts, n_experts_per_token, total_top_k, bs.
# > 
# > Then, the fused_moe kernel flow is:
# > (1) Quant activations to MXFP4: aiter per-1x32 dynamic quantization of hidden_states.
# > (2) Stage 1 GEMM + activation (per token i, per assigned expert j):
# >     - gate = x_i @ W_gate_j.T          # [d_hidden] x [d_expert, d_hidden].T -> [d_expert]
# >     - up   = x_i @ W_up_j.T            # [d_hidden] x [d_expert, d_hidden].T -> [d_expert]
# >     - intermediate = SiLU(gate) * up    # SwiGLU activation, -> [d_expert]
# >     (W_gate and W_up are fused as gate_up_weight, so this is one a4w4 GEMM + fused activation)
# > (3) Stage 2 GEMM:
# >     - expert_out = intermediate @ W_down_j.T  # [d_expert] x [d_hidden, d_expert].T -> [d_hidden]
# > (4) Weighted reduction:
# >     - output_i += w_ij * expert_out     # accumulate across top_k experts
# > All weight GEMMs are a4w4 (MXFP4 activations x MXFP4 weights, per-1x32 block scaling).
# > The AITER CK kernel fuses all of the above into a 2-stage pipeline across all tokens and experts.
# > 
# > DeepSeek-R1 MoE specs:
# >   - hidden_size = 7168, moe_intermediate_size = 2048
# >   - 256 routed experts + 1 shared expert (total 257), top-8 routed + 1 shared = 9 per token
# >   - 58 MoE layers (layer 3-60)
# >   - The shared expert processes ALL tokens unconditionally (weight=1.0)
# > 
# > d_hidden_pad and d_expert_pad are the dimensions padded to 256-alignment for the CK kernel.
# > 
# > The ranking criteria is the geometric mean of the benchmark results.
# > 
# > ```
# > The AITER reference performance is (E includes shared expert, top_k = routed + shared):
# >   bs     E  d_hidden  d_expert  top_k  time[us]
# >   16   257      7168       256      9    152.7
# >  128   257      7168       256      9    239.0
# >  512   257      7168       256      9    336.5
# >   16    33      7168       512      9    106.2
# >  128    33      7168       512      9    141.1
# >  512    33      7168       512      9    225.0
# >  512    33      7168      2048      9    380.4
# > ```
# > 
# > Input:
# >   - hidden_states:                  [M, d_hidden]                          bf16
# >   - gate_up_weight:                 [E, 2*d_expert_pad, d_hidden_pad//2]   fp4x2 (raw, before shuffle)
# >   - down_weight:                    [E, d_hidden_pad, d_expert_pad//2]     fp4x2 (raw, before shuffle)
# >   - gate_up_weight_scale:           [E, 2*d_expert_pad, d_hidden_pad//32]  e8m0  (raw, before shuffle)
# >   - down_weight_scale:              [E, d_hidden_pad, d_expert_pad//32]    e8m0  (raw, before shuffle)
# >   - gate_up_weight_shuffled:        [E, 2*d_expert_pad, d_hidden_pad//2]   fp4x2 (pre-shuffled for CK)
# >   - down_weight_shuffled:           [E, d_hidden_pad, d_expert_pad//2]     fp4x2 (pre-shuffled for CK)
# >   - gate_up_weight_scale_shuffled:  [padded, flat]                         e8m0  (pre-shuffled for CK)
# >   - down_weight_scale_shuffled:     [padded, flat]                         e8m0  (pre-shuffled for CK)
# >   - topk_weights:                   [M, total_top_k]                       float32
# >   - topk_ids:                       [M, total_top_k]                       int32
# >   - config:                         dict with dimensions
# > 
# > Output:
# >   - output: [M, d_hidden] bf16
# The deadline for this leaderboard is 2026-04-07 07:59:00+00:00

# You can automatically route this file to specific GPUs by adding a line
# `#!POPCORN gpus <GPUs>` to the header of this file.
# Happy hacking!

import torch
from typing import Dict
from task import input_t, output_t

from aiter import ActivationType, QuantType
from aiter.fused_moe import fused_moe


def custom_kernel(data: input_t) -> output_t:
    """
    Submission template for DeepSeek-R1 MXFP4 MoE kernel.

    Input data tuple:
        hidden_states:                [M, d_hidden]                           bf16
        gate_up_weight:               [E, 2*d_expert_pad, d_hidden_pad//2]    fp4x2  (raw)
        down_weight:                  [E, d_hidden_pad, d_expert_pad//2]      fp4x2  (raw)
        gate_up_weight_scale:         [E, 2*d_expert_pad, scale_K]            e8m0   (raw)
        down_weight_scale:            [E, d_hidden_pad, scale_K]              e8m0   (raw)
        gate_up_weight_shuffled:      [E, 2*d_expert_pad, d_hidden_pad//2]    fp4x2  (shuffled)
        down_weight_shuffled:         [E, d_hidden_pad, d_expert_pad//2]      fp4x2  (shuffled)
        gate_up_weight_scale_shuffled:[padded, flat]                          e8m0   (shuffled)
        down_weight_scale_shuffled:   [padded, flat]                          e8m0   (shuffled)
        topk_weights:                 [M, total_top_k]                        float32
        topk_ids:                     [M, total_top_k]                        int32
        config:                       dict

    Returns:
        output: [M, d_hidden] bf16
    """
    (
        hidden_states,
        gate_up_weight,
        down_weight,
        gate_up_weight_scale,
        down_weight_scale,
        gate_up_weight_shuffled,
        down_weight_shuffled,
        gate_up_weight_scale_shuffled,
        down_weight_scale_shuffled,
        topk_weights,
        topk_ids,
        config,
    ) = data

    hidden_pad = config["d_hidden_pad"] - config["d_hidden"]
    intermediate_pad = config["d_expert_pad"] - config["d_expert"]

    output = fused_moe(
        hidden_states,
        gate_up_weight_shuffled,
        down_weight_shuffled,
        topk_weights,
        topk_ids,
        expert_mask=None,
        activation=ActivationType.Silu,
        quant_type=QuantType.per_1x32,
        doweight_stage1=False,
        w1_scale=gate_up_weight_scale_shuffled,
        w2_scale=down_weight_scale_shuffled,
        a1_scale=None,
        a2_scale=None,
        hidden_pad=hidden_pad,
        intermediate_pad=intermediate_pad,
    )

    return output

scrolls · 163 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON