Skip to content
KernelIndex
Search⌘K

submission 564613

Infatoshi · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 50 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-moe-mxfp4-564613?include=source"
interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, fp32, fp8_e8m0, int32, mxfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
AMD MXFP4 MoEsuite of 7 cases
AMD Instinct MI355X
185.7µs
#658 of 782
2026-03-16

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:4e81525c2ef26983989262d3d5b6594ee4a6249c70d59f5ae449e7ae06a8690e
license declaredunknown
license concludedunknown
authorsInfatoshi
imported2026-08-26

Kernel source

submission.py50 lines
#!POPCORN leaderboard amd-moe-mxfp4
#!POPCORN gpu MI355X
# AGENT_LOOP_META: {"attempt": 40, "generator": {"kind": "llm", "model": "(codex default)", "parallel_agents": 3, "provider": "codex_cli", "use_plan": true}, "gpu": "MI355X", "leaderboard": "amd-moe-mxfp4", "policy_profile": {"family": "kernel_explore", "focus": "stabilize routing and shuffled-weight semantics", "name": "contract_repair", "trigger_signals": ["contract_repair", "runtime_repair", "submission_repair"]}, "problem": "moe_mxfp4", "variant": {"BLOCK_SIZE": 256, "NUM_WARPS": 4, "SORT_BY_EXPERT": true, "family": "kernel_explore", "strategy": "routing_prototype", "variant_name": "routing_swiglu_256"}, "variant_index": 2}

from aiter import ActivationType, QuantType
from aiter.fused_moe import fused_moe
from task import input_t, output_t

_ACTIVATION = ActivationType.Silu
_QUANT_TYPE = QuantType.per_1x32
_FUSED_MOE = fused_moe


def custom_kernel(data: input_t) -> output_t:
    (
        hidden_states,
        _,
        _,
        _,
        _,
        gate_up_weight_shuffled,
        down_weight_shuffled,
        gate_up_weight_scale_shuffled,
        down_weight_scale_shuffled,
        topk_weights,
        topk_ids,
        config,
    ) = data

    hidden_pad = down_weight_shuffled.shape[1] - hidden_states.shape[1]
    intermediate_pad = (gate_up_weight_shuffled.shape[1] >> 1) - config["d_expert"]

    return _FUSED_MOE(
        hidden_states,
        gate_up_weight_shuffled,
        down_weight_shuffled,
        topk_weights,
        topk_ids,
        expert_mask=None,
        activation=_ACTIVATION,
        quant_type=_QUANT_TYPE,
        doweight_stage1=False,
        w1_scale=gate_up_weight_scale_shuffled,
        w2_scale=down_weight_scale_shuffled,
        a1_scale=None,
        a2_scale=None,
        hidden_pad=hidden_pad,
        intermediate_pad=intermediate_pad,
    )
scrolls · 50 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON