Skip to content
KernelIndex
Search⌘K

submission 718275

jin051623 · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 66 lines, June 9 Researcher Reciprocity License v1.0.

submission1.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-amd-mxfp4-mm-718275?include=source"
interfacepython
Compatibility
measured onAMD Instinct MI355X
declared hardwareAMD Instinct MI355X
architecturesgfx950
dtypesbf16, mxfp4

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
AMD MXFP4 GEMMsuite of 6 cases
AMD Instinct MI355X
24.1µs
#941 of 1143
2026-04-04

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:256d1d12bc84f795b11aa22ce06bd3ce5b5c3498c125d8d9e1057fb0afd70a44
license declaredunknown
license concludedunknown
authorsjin051623
imported2026-08-26

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

fp4FP4 quant + FP4 GEMM reference: bf16 A, MXFP4 B -> MXFP4 per-1x32 quant A -> gemm_a4w4 -> bf16 C.

Kernel source

submission1.py66 lines
#!POPCORN leaderboard amd-mxfp4-mm
#!POPCORN gpu MI355X

"""
Optimized submission: same semantics as reference but lower Python overhead.
FP4 quant + FP4 GEMM reference: bf16 A, MXFP4 B -> MXFP4 per-1x32 quant A -> gemm_a4w4 -> bf16 C.
"""

from task import input_t, output_t

# Module-scope imports to avoid per-call overhead
import aiter
from aiter import dtypes
from aiter.ops.triton.quant import dynamic_mxfp4_quant  # patched kernel per reference
from aiter.utility.fp4_utils import e8m0_shuffle

# Keep SCALE_GROUP_SIZE visible if needed by future local checks
SCALE_GROUP_SIZE = 32


# Module-scope helper: quantizes bf16 -> packed fp4x2 + shuffled E8M0 scales
def _quant_mxfp4_per_1x32(x, shuffle: bool = True):
    """
    x: bf16 [M, K]
    returns:
      - x_fp4_packed: dtypes.fp4x2 view (packed)
      - bs_e8m0_shuffled: dtypes.fp8_e8m0 view (shuffled if requested)
    """
    # dynamic_mxfp4_quant is the patched, Triton-backed quantizer referenced in the tests
    x_fp4, bs_e8m0 = dynamic_mxfp4_quant(x)
    if shuffle:
        bs_e8m0 = e8m0_shuffle(bs_e8m0)
    # Return views in the exact dtype shapes expected by gemm_a4w4
    return x_fp4.view(dtypes.fp4x2), bs_e8m0.view(dtypes.fp8_e8m0)


def custom_kernel(data: input_t) -> output_t:
    """
    Optimized hot-path:
      - Quantize A (bf16) to MXFP4 per-1x32 (packed)
      - Call aiter.gemm_a4w4 with pre-shuffled B and scales
      - Return bf16 C
    """
    # Unpack only what we need; keep variable names consistent with reference harness
    A, B, B_q, B_shuffle, B_scale_sh = data

    # Ensure A is contiguous once (avoid unconditional copy if already contiguous)
    if not A.is_contiguous():
        A = A.contiguous()

    # Quantize A to MXFP4 per-1x32 and get shuffled scales
    A_q, A_scale_sh = _quant_mxfp4_per_1x32(A, shuffle=True)

    # Call the Triton-backed GEMM (keeps correctness and uses optimized backend)
    # bpreshuffle=True matches the reference harness expectations
    out = aiter.gemm_a4w4(
        A_q,
        B_shuffle,
        A_scale_sh,
        B_scale_sh,
        dtype=dtypes.bf16,
        bpreshuffle=True,
    )

    return out
scrolls · 66 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON