Skip to content
KernelIndex
Search⌘K

submission 715843

Elán Zainos Corona · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 50 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-moe-mxfp4-715843?include=source"
interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, fp32, fp8_e8m0, int32, mxfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
AMD MXFP4 MoEsuite of 7 cases
AMD Instinct MI355X
177.8µs
#387 of 782
2026-04-03

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:b3760201e8890096c97a81c6e654d8233b74959f463ed6f4e7b6416217f4f564
license declaredunknown
license concludedunknown
authorsElán Zainos Corona
imported2026-08-26

Kernel source

submission.py50 lines
# -*- coding: utf-8 -*-
import torch
import os

def custom_kernel(data):
    """
    V35.0: Wrapper-Restoration Protocol.
    Reversion a contenedor heuristico para recuperacion de resolucion de nucleos C++.
    """
    # Registro de ejecucion obligatorio
    if not hasattr(custom_kernel, "_log_init"):
        os.write(1, b"[SENTINEL_LOG] V35.0 Inicializada. Reversion a despachador heuristico de alto nivel.\n")
        custom_kernel._log_init = True

    # 1. Extraccion de Payload
    is_t = type(data) is tuple
    h = data[0] if is_t else data.hidden_states
    w1 = data[5] if is_t else data.gate_up_weight_shuffled
    w2 = data[6] if is_t else data.down_weight_shuffled
    s1 = data[7] if is_t else data.gate_up_weight_scale_shuffled
    s2 = data[8] if is_t else data.down_weight_scale_shuffled
    tw = data[9] if is_t else data.topk_weights
    ti = data[10] if is_t else data.topk_ids
    c = data[11] if is_t else data.config

    # 2. Importacion Diferida (Lazy Loading)
    try:
        from aiter.fused_moe import fused_moe
        from aiter import ActivationType, QuantType
    except ImportError:
        os.write(1, b"[SENTINEL_ERROR] Falla de importacion de libreria externa.\n")
        return torch.zeros_like(h).to(torch.bfloat16)

    # 3. Calculo de Alineacion de Memoria
    h_pad = c.get("d_hidden_pad", c["d_hidden"]) - c["d_hidden"]
    i_pad = c.get("d_expert_pad", c["d_expert"]) - c["d_expert"]

    # 4. Despacho de Inferencia (Wrapper Reintegrado)
    output = fused_moe(
        h, w1, w2, tw, ti,
        activation=ActivationType.Silu,
        quant_type=QuantType.per_1x32,
        w1_scale=s1,
        w2_scale=s2,
        hidden_pad=h_pad,
        intermediate_pad=i_pad,
        doweight_stage1=False 
    )

    return output.to(torch.bfloat16)
scrolls · 50 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON