Skip to content
KernelIndex
Search⌘K

submission 678801

jiulvke · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 44 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-mxfp4-mm-678801?include=source"
interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, mxfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
AMD MXFP4 GEMMsuite of 6 cases
AMD Instinct MI355X
23.7µs
#801 of 1143
2026-03-31

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:7e8414ee33dc774fe884b9fead03069d07824c6c28740b60dae77345ed95c3a3
license declaredunknown
license concludedunknown
authorsjiulvke
imported2026-08-26

Kernel source

submission.py44 lines
import torch
import aiter

# --- 战前准备:将所有“重武器”提前搬出库房,消除函数内查找开销 ---
# 1. 预存底层 C++ 算子句柄 (跳过 aiter 的 Python 封装)
_Q_OP = aiter.ops.triton.quant.dynamic_mxfp4_quant
_S_OP = aiter.utility.fp4_utils.e8m0_shuffle
_G_OP = torch.ops.aiter.gemm_a4w4

# 2. 预存数据类型 (避免在循环中访问 aiter.dtypes)
_FP4_V = aiter.dtypes.fp4x2
_E8M0_V = aiter.dtypes.fp8_e8m0
_BF16_TYPE = torch.bfloat16

# 3. 预存 Bfloat16 的枚举值,压榨参数解析速度
# 在底层 ScalarType 中,bfloat16 通常对应 15
_DTYPE_ENUM = 15 

@torch.inference_mode()
def custom_kernel(data):
    """
    MM 赛道极致方案:
    1. 使用下标访问 data,避开 tuple 解包和命名查找。
    2. 使用预存的底层 Op 直接派发,跳过 Python 层逻辑。
    3. 彻底消除显存申请。
    """
    # 极速量化 A
    # aq: [M, K/2], as_: [M, K/32]
    aq, as_ = _Q_OP(data[0])
    
    # 终极一击:直接调用 C++ Schema 接口
    # Schema: (A, B, A_scale, B_scale, bias, dtype, alpha, beta, bpreshuffle)
    # 通过 view(_FP4_V) 实现零拷贝类型转换
    return _G_OP(
        aq.view(_FP4_V),     # A: data[0] 的量化版
        data[3],             # B_sh: data[3]
        _S_OP(as_).view(_E8M0_V), # A_scale_sh
        data[4],             # B_scale_sh: data[4]
        None,                # bias
        _BF16_TYPE,          # dtype
        1.0,                 # alpha
        0.0,                 # beta
        True                 # bpreshuffle
    )
scrolls · 44 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON