Skip to content
KernelIndex
Search⌘K

submission 682356

tiendreiliass0x · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 35 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-moe-mxfp4-682356?include=source"
interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, fp32, fp8_e8m0, int32, mxfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
AMD MXFP4 MoEsuite of 7 cases
AMD Instinct MI355X
177.4µs
#370 of 782
2026-03-31

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:45c0ff82bb95ca3d0d7dfc2bc958b8867d9829bf5723ef137712128a67a67474
license declaredunknown
license concludedunknown
authorstiendreiliass0x
imported2026-08-26

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

fp4"""MXFP4 MoE with aiter.fused_moe using proper enum types."""

Kernel source

submission.py35 lines
#!POPCORN leaderboard amd-moe-mxfp4
#!POPCORN gpu MI355X

"""MXFP4 MoE with aiter.fused_moe using proper enum types."""

import torch
from task import input_t, output_t
from aiter.fused_moe import fused_moe
from aiter import ActivationType, QuantType


def custom_kernel(data: input_t) -> output_t:
    x = data[0]
    w1_shuffled = data[5]
    w2_shuffled = data[6]
    w1_shuffled_scale = data[7]
    w2_shuffled_scale = data[8]
    routing_weights = data[9]
    routing_indices = data[10]

    out = fused_moe(
        hidden_states=x,
        w1=w1_shuffled,
        w2=w2_shuffled,
        topk_weight=routing_weights,
        topk_ids=routing_indices,
        expert_mask=None,
        activation=ActivationType.Silu,
        quant_type=QuantType.per_1x32,
        doweight_stage1=False,
        w1_scale=w1_shuffled_scale,
        w2_scale=w2_shuffled_scale,
    )
    return out
scrolls · 35 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON