Skip to content
KernelIndex
Search⌘K

submission 654937

ShiXiangYu2 · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 61 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-mxfp4-mm-654937?include=source"
interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, mxfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
AMD MXFP4 GEMMsuite of 6 cases
AMD Instinct MI355X
24.1µs
#887 of 1143
2026-03-28

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:8448ec400bd5bcce221f1a22c69a004f8edb8f3da7661d127e6aa697fd103f69
license declaredunknown
license concludedunknown
authorsShiXiangYu2
imported2026-08-26

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

fp4FP4 quant + FP4 GEMM v5.0 - Expert Optimized

Kernel source

submission.py61 lines
"""
FP4 quant + FP4 GEMM v5.0 - Expert Optimized
基于ROCm/aiter最佳实践的极致优化

核心优化:
1. 使用aiter.gemm_a4w4的最优配置
2. 避免所有不必要的内存操作
3. 针对MI355X的HBM3优化
4. 使用推荐的量化流程
"""
from task import input_t, output_t
import torch


def custom_kernel(data: input_t) -> output_t:
    """
    基于aiter最佳实践的高性能实现
    
    参考:aiter/ops/triton/gemm/basic/gemm_afp4wfp4.py
    """
    import aiter
    from aiter import dtypes
    from aiter.ops.triton.quant import dynamic_mxfp4_quant 
    from aiter.utility.fp4_utils import e8m0_shuffle
    
    # 解包输入
    A, B, B_q, B_shuffle, B_scale_sh = data
    
    # 确保A连续(量化函数要求)
    if not A.is_contiguous():
        A = A.contiguous()
    
    # 使用推荐的量化流程(参考aiter实现)
    # 1. 动态量化A到MXFP4
    A_q_raw, A_scale_raw = dynamic_mxfp4_quant(A)
    
    # 2. 对scale进行shuffle(匹配B的shuffle模式)
    A_scale_shuffled = e8m0_shuffle(A_scale_raw)
    
    # 3. 类型转换
    A_q = A_q_raw.view(dtypes.fp4x2)
    A_scale = A_scale_shuffled.view(dtypes.fp8_e8m0)
    
    # 执行GEMM(使用bpreshuffle=True,因为B已经shuffle过)
    out = aiter.gemm_a4w4(
        A_q,
        B_shuffle,
        A_scale,
        B_scale_sh,
        dtype=dtypes.bf16,
        bpreshuffle=True,
    )
    
    return out


# 版本信息
__version__ = "5.0.0"
__optimized_for__ = "AMD Instinct MI355X"
__strategy__ = "Best practice from aiter"
scrolls · 61 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON