submission 678801
jiulvke · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 44 lines, June 9 Researcher Reciprocity License v1.0.
submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-mxfp4-mm-678801?include=source"interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, mxfp4
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:7e8414ee33dc774fe884b9fead03069d07824c6c28740b60dae77345ed95c3a3
license declaredunknown
license concludedunknown
authorsjiulvke
imported2026-08-26
Kernel source
submission.py44 lines
import torch
import aiter
# --- 战前准备:将所有“重武器”提前搬出库房,消除函数内查找开销 ---
# 1. 预存底层 C++ 算子句柄 (跳过 aiter 的 Python 封装)
_Q_OP = aiter.ops.triton.quant.dynamic_mxfp4_quant
_S_OP = aiter.utility.fp4_utils.e8m0_shuffle
_G_OP = torch.ops.aiter.gemm_a4w4
# 2. 预存数据类型 (避免在循环中访问 aiter.dtypes)
_FP4_V = aiter.dtypes.fp4x2
_E8M0_V = aiter.dtypes.fp8_e8m0
_BF16_TYPE = torch.bfloat16
# 3. 预存 Bfloat16 的枚举值,压榨参数解析速度
# 在底层 ScalarType 中,bfloat16 通常对应 15
_DTYPE_ENUM = 15
@torch.inference_mode()
def custom_kernel(data):
"""
MM 赛道极致方案:
1. 使用下标访问 data,避开 tuple 解包和命名查找。
2. 使用预存的底层 Op 直接派发,跳过 Python 层逻辑。
3. 彻底消除显存申请。
"""
# 极速量化 A
# aq: [M, K/2], as_: [M, K/32]
aq, as_ = _Q_OP(data[0])
# 终极一击:直接调用 C++ Schema 接口
# Schema: (A, B, A_scale, B_scale, bias, dtype, alpha, beta, bpreshuffle)
# 通过 view(_FP4_V) 实现零拷贝类型转换
return _G_OP(
aq.view(_FP4_V), # A: data[0] 的量化版
data[3], # B_sh: data[3]
_S_OP(as_).view(_E8M0_V), # A_scale_sh
data[4], # B_scale_sh: data[4]
None, # bias
_BF16_TYPE, # dtype
1.0, # alpha
0.0, # beta
True # bpreshuffle
)scrolls · 44 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON