Skip to content
KernelIndex
Search⌘K

submission 583919

Zobin Huang · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 70 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-moe-mxfp4-583919?include=source"
interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, fp32, fp8_e8m0, int32, mxfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
AMD MXFP4 MoEsuite of 7 cases
AMD Instinct MI355X
188.6µs
#740 of 782
2026-03-18

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:8d124920c2a49cde0f5c8019c3d56f357c1871d4178d0b37d29bea6956d0c7ab
license declaredunknown
license concludedunknown
authorsZobin Huang
imported2026-08-26

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

fp4"""Optimized MoE MXFP4 submission for AMD Instinct MI355X.

Kernel source

submission.py70 lines
"""Optimized MoE MXFP4 submission for AMD Instinct MI355X.

Uses AITER's fused_moe with per-config block_size_M tuning.
The block_size_M controls the tile size for the CK/ASM kernel dispatch:
- For many-expert configs (E=257): smaller block_m=32 reduces padding waste
- For fewer-expert configs (E=33): larger block_m balances CU utilization
"""
from task import input_t, output_t
import torch

from aiter import ActivationType, QuantType
from aiter.fused_moe import fused_moe


def custom_kernel(data: input_t) -> output_t:
    (
        hidden_states,
        gate_up_weight,
        down_weight,
        gate_up_weight_scale,
        down_weight_scale,
        gate_up_weight_shuffled,
        down_weight_shuffled,
        gate_up_weight_scale_shuffled,
        down_weight_scale_shuffled,
        topk_weights,
        topk_ids,
        config,
    ) = data

    hidden_pad = config["d_hidden_pad"] - config["d_hidden"]
    intermediate_pad = config["d_expert_pad"] - config["d_expert"]

    # Select block_size_M based on problem shape
    E_total = config["n_routed_experts"] + config["n_shared_experts"]
    bs = config["bs"]
    d_expert = config["d_expert"]

    # For many-expert configs, block_m=32 minimizes padding waste
    # For fewer-expert configs, heuristic auto-selection is usually good
    if E_total > 64:
        block_m = 32
    elif d_expert >= 2048:
        block_m = 64
    elif bs >= 128:
        block_m = 128
    else:
        block_m = 64

    output = fused_moe(
        hidden_states,
        gate_up_weight_shuffled,
        down_weight_shuffled,
        topk_weights,
        topk_ids,
        expert_mask=None,
        activation=ActivationType.Silu,
        quant_type=QuantType.per_1x32,
        doweight_stage1=False,
        w1_scale=gate_up_weight_scale_shuffled,
        w2_scale=down_weight_scale_shuffled,
        a1_scale=None,
        a2_scale=None,
        block_size_M=block_m,
        hidden_pad=hidden_pad,
        intermediate_pad=intermediate_pad,
    )

    return output
scrolls · 70 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON