Skip to content
KernelIndex
Search⌘K

submission 590226

leaatimberini · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 42 lines, June 9 Researcher Reciprocity License v1.0.

submission_moe.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-moe-mxfp4-590226?include=source"
interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, fp32, fp8_e8m0, int32, mxfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
AMD MXFP4 MoEsuite of 7 cases
AMD Instinct MI355X
180.5µs
#490 of 782
2026-03-19

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:9cc1d55fe22d455700e00d134076b7d518d72524d70f00dc1a1450b40a3bbb40
license declaredunknown
license concludedunknown
authorsleaatimberini
imported2026-08-26

Kernel source

submission_moe.py42 lines
#!POPCORN leaderboard amd-moe-mxfp4
#!POPCORN gpu MI355X

import torch
import aiter
from aiter import ActivationType, QuantType
from utils import make_match_reference

class MoE_V1_Dispatcher:
    def __init__(self):
        self.hidden_pad = None
        self.intermediate_pad = None

    @torch.compiler.disable
    def __call__(self, data):
        # Despacho limpio y directo sin overhead extra
        if self.hidden_pad is None:
            config = data[11]
            self.hidden_pad = config["d_hidden_pad"] - config["d_hidden"]
            self.intermediate_pad = config["d_expert_pad"] - config["d_expert"]

        return aiter.fused_moe.fused_moe(
            data[0],   
            data[5],   
            data[6],   
            data[9],   
            data[10],  
            expert_mask=None,
            activation=ActivationType.Silu,
            quant_type=QuantType.per_1x32, 
            doweight_stage1=False,  # Mantenemos las heurísticas seguras de AMD
            w1_scale=data[7],  
            w2_scale=data[8],  
            a1_scale=None,
            a2_scale=None,
            hidden_pad=self.hidden_pad,
            intermediate_pad=self.intermediate_pad,
        )

custom_kernel = MoE_V1_Dispatcher()
submission = custom_kernel
check_implementation = make_match_reference(custom_kernel, rtol=2e-2, atol=2e-2)
scrolls · 42 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON