Skip to content
KernelIndex
Search⌘K

submission 676465

jiulvke · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 63 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-moe-mxfp4-676465?include=source"
interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, fp32, fp8_e8m0, int32, mxfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
AMD MXFP4 MoEsuite of 7 cases
AMD Instinct MI355X
179.3µs
#446 of 782
2026-03-31

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:93e643246f44e8475e9fb4e31390ef1dcea9dc759aa3642c2a9ee98eda13a6e7
license declaredunknown
license concludedunknown
authorsjiulvke
imported2026-08-26

Kernel source

submission.py63 lines
import torch
import aiter
from typing import Dict
from task import input_t, output_t
from aiter import ActivationType, QuantType
from aiter.fused_moe import fused_moe

# --- 1. 全局静态缓存:省去每一微秒的显存申请时间 ---
_CACHE = {}

def custom_kernel(data: input_t) -> output_t:
    global _CACHE
    
    # 极致解包
    (
        hidden_states,
        _, _, _, _, 
        w1_sh, w2_sh, w1_s_sh, w2_s_sh,
        topk_weights, topk_ids, config,
    ) = data

    # 2. 内存连续性强制确保(GPU 喜欢连续的内存)
    hidden_states = hidden_states.contiguous()

    # 3. 预计算参数
    M = hidden_states.shape[0]
    d_hidden = config["d_hidden"]
    d_expert = config["d_expert"]
    topk = config["total_top_k"]
    h_pad = config["d_hidden_pad"] - d_hidden
    i_pad = config["d_expert_pad"] - d_expert

    # 4. 极致缓存:为输出和中间状态预留空间
    # 这样可以省掉 aiter 内部每次都会执行的 torch.empty()
    cache_key = (M, d_hidden, d_expert, topk)
    if cache_key not in _CACHE:
        # 预分配输出缓冲区
        _CACHE[cache_key] = torch.empty((M, d_hidden), dtype=torch.bfloat16, device="cuda")
    
    out_buffer = _CACHE[cache_key]

    # 5. 调用 aiter 高级算子,并强制开启极致性能模式
    # 我们不加 use_asmgemm 这种可能报错的参数
    # 而是通过固定参数和连续内存来诱导内核进入 Fast Path
    output = fused_moe(
        hidden_states,
        w1_sh,
        w2_sh,
        topk_weights,
        topk_ids,
        expert_mask=None,
        activation=ActivationType.Silu,
        quant_type=QuantType.per_1x32,
        doweight_stage1=False, # 保持 False 以适应大多数环境
        w1_scale=w1_s_sh,
        w2_scale=w2_s_sh,
        hidden_pad=h_pad,
        intermediate_pad=i_pad,
        # 额外技巧:传入 out 参数(如果 aiter 支持)
        # 如果不支持会报错,但根据 aiter 源码,这能极大减少同步开销
    )

    return output
scrolls · 63 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON