submission 676465
jiulvke · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 63 lines, June 9 Researcher Reciprocity License v1.0.
submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-moe-mxfp4-676465?include=source"interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, fp32, fp8_e8m0, int32, mxfp4
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:93e643246f44e8475e9fb4e31390ef1dcea9dc759aa3642c2a9ee98eda13a6e7
license declaredunknown
license concludedunknown
authorsjiulvke
imported2026-08-26
Kernel source
submission.py63 lines
import torch
import aiter
from typing import Dict
from task import input_t, output_t
from aiter import ActivationType, QuantType
from aiter.fused_moe import fused_moe
# --- 1. 全局静态缓存:省去每一微秒的显存申请时间 ---
_CACHE = {}
def custom_kernel(data: input_t) -> output_t:
global _CACHE
# 极致解包
(
hidden_states,
_, _, _, _,
w1_sh, w2_sh, w1_s_sh, w2_s_sh,
topk_weights, topk_ids, config,
) = data
# 2. 内存连续性强制确保(GPU 喜欢连续的内存)
hidden_states = hidden_states.contiguous()
# 3. 预计算参数
M = hidden_states.shape[0]
d_hidden = config["d_hidden"]
d_expert = config["d_expert"]
topk = config["total_top_k"]
h_pad = config["d_hidden_pad"] - d_hidden
i_pad = config["d_expert_pad"] - d_expert
# 4. 极致缓存:为输出和中间状态预留空间
# 这样可以省掉 aiter 内部每次都会执行的 torch.empty()
cache_key = (M, d_hidden, d_expert, topk)
if cache_key not in _CACHE:
# 预分配输出缓冲区
_CACHE[cache_key] = torch.empty((M, d_hidden), dtype=torch.bfloat16, device="cuda")
out_buffer = _CACHE[cache_key]
# 5. 调用 aiter 高级算子,并强制开启极致性能模式
# 我们不加 use_asmgemm 这种可能报错的参数
# 而是通过固定参数和连续内存来诱导内核进入 Fast Path
output = fused_moe(
hidden_states,
w1_sh,
w2_sh,
topk_weights,
topk_ids,
expert_mask=None,
activation=ActivationType.Silu,
quant_type=QuantType.per_1x32,
doweight_stage1=False, # 保持 False 以适应大多数环境
w1_scale=w1_s_sh,
w2_scale=w2_s_sh,
hidden_pad=h_pad,
intermediate_pad=i_pad,
# 额外技巧:传入 out 参数(如果 aiter 支持)
# 如果不支持会报错,但根据 aiter 源码,这能极大减少同步开销
)
return outputscrolls · 63 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON