Skip to content
KernelIndex
Search⌘K

submission 694654

DiegoCao · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 63 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-moe-mxfp4-694654?include=source"
interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, fp32, fp8_e8m0, int32, mxfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
AMD MXFP4 MoEsuite of 7 cases
AMD Instinct MI355X
170.0µs
#314 of 782
2026-04-02

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:73fe1f7e737fefbfbab4829c943989112991757f64ff709c2f916e2764020b0f
license declaredunknown
license concludedunknown
authorsDiegoCao
imported2026-08-15

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

fp4MoE MXFP4 v13: Adaptive splitk via env var manipulation + cache clearing.
split-kMoE MXFP4 v13: Adaptive splitk via env var manipulation + cache clearing.

Kernel source

submission.py63 lines
#!POPCORN leaderboard amd-moe-mxfp4
#!POPCORN gpu MI355X

"""
MoE MXFP4 v13: Adaptive splitk via env var manipulation + cache clearing.

The fused_moe's get_ksplit uses @lru_cache keyed on (token, topk, expert, inter_dim, model_dim).
Since different workloads have different (token, expert, inter_dim), each gets its own
cache entry. The AITER_KSPLIT env var is read INSIDE get_ksplit on cache miss.

Strategy: set AITER_KSPLIT=2 for low-occupancy workloads, 0 for others.
Since each workload size is unique, the cache will miss on first call for each size.
We just need to set the env var BEFORE the first call to fused_moe for each workload.
"""

import os
import torch
from task import input_t, output_t

os.environ.setdefault("AITER_USE_OPUS_MOE_SORTING", "1")

from aiter import ActivationType, QuantType
from aiter.fused_moe import fused_moe

_SILU = ActivationType.Silu
_PER_1x32 = QuantType.per_1x32
_fused_moe = fused_moe


def custom_kernel(data: input_t) -> output_t:
    (
        hidden_states,
        _, _, _, _,
        w1, w2,
        w1_scale, w2_scale,
        topk_weights,
        topk_ids,
        config,
    ) = data

    M = hidden_states.shape[0]
    topk = topk_ids.shape[1]
    E = w1.shape[0]

    # Adaptive splitk: set env var before fused_moe's lru_cache lookup
    # Each unique (M, E, d_expert) will be a cache miss on first call
    tokens_per_expert = M * topk / E
    if tokens_per_expert < 64 and M <= 128:
        os.environ["AITER_KSPLIT"] = "2"
    else:
        os.environ["AITER_KSPLIT"] = "0"

    return _fused_moe(
        hidden_states, w1, w2,
        topk_weights, topk_ids,
        activation=_SILU,
        quant_type=_PER_1x32,
        w1_scale=w1_scale,
        w2_scale=w2_scale,
        hidden_pad=config["d_hidden_pad"] - config["d_hidden"],
        intermediate_pad=config["d_expert_pad"] - config["d_expert"],
    )
scrolls · 63 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON