submission 654937
ShiXiangYu2 · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 61 lines, June 9 Researcher Reciprocity License v1.0.
submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-mxfp4-mm-654937?include=source"interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, mxfp4
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:8448ec400bd5bcce221f1a22c69a004f8edb8f3da7661d127e6aa697fd103f69
license declaredunknown
license concludedunknown
authorsShiXiangYu2
imported2026-08-26
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
fp4
FP4 quant + FP4 GEMM v5.0 - Expert OptimizedKernel source
submission.py61 lines
"""
FP4 quant + FP4 GEMM v5.0 - Expert Optimized
基于ROCm/aiter最佳实践的极致优化
核心优化:
1. 使用aiter.gemm_a4w4的最优配置
2. 避免所有不必要的内存操作
3. 针对MI355X的HBM3优化
4. 使用推荐的量化流程
"""
from task import input_t, output_t
import torch
def custom_kernel(data: input_t) -> output_t:
"""
基于aiter最佳实践的高性能实现
参考:aiter/ops/triton/gemm/basic/gemm_afp4wfp4.py
"""
import aiter
from aiter import dtypes
from aiter.ops.triton.quant import dynamic_mxfp4_quant
from aiter.utility.fp4_utils import e8m0_shuffle
# 解包输入
A, B, B_q, B_shuffle, B_scale_sh = data
# 确保A连续(量化函数要求)
if not A.is_contiguous():
A = A.contiguous()
# 使用推荐的量化流程(参考aiter实现)
# 1. 动态量化A到MXFP4
A_q_raw, A_scale_raw = dynamic_mxfp4_quant(A)
# 2. 对scale进行shuffle(匹配B的shuffle模式)
A_scale_shuffled = e8m0_shuffle(A_scale_raw)
# 3. 类型转换
A_q = A_q_raw.view(dtypes.fp4x2)
A_scale = A_scale_shuffled.view(dtypes.fp8_e8m0)
# 执行GEMM(使用bpreshuffle=True,因为B已经shuffle过)
out = aiter.gemm_a4w4(
A_q,
B_shuffle,
A_scale,
B_scale_sh,
dtype=dtypes.bf16,
bpreshuffle=True,
)
return out
# 版本信息
__version__ = "5.0.0"
__optimized_for__ = "AMD Instinct MI355X"
__strategy__ = "Best practice from aiter"
scrolls · 61 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON