Skip to content
KernelIndex
Search⌘K

submission 776985

ddwinterdd · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 82 lines, June 9 Researcher Reciprocity License v1.0.

submission_H100.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-matmul-v2-776985?include=source"
interfacepython
Compatibility
measured onNVIDIA H100
declared hardwareNVIDIA H100
architecturessm_90
dtypesfp16

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
FP16 matmulsuite of 8 cases
NVIDIA H100
246.5µs
#11 of 28
2026-04-20

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:3157a45682b1fe4f844cf12ac4f295932813a070d6390b9a887537100f6d4999
license declaredunknown
license concludedunknown
authorsddwinterdd
imported2026-08-15

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

autotune@triton.autotune(
mmaaccu = tl.dot(a_block, b_block, acc=accu)
num-warps = 8triton.Config({}, num_warps=8),
tile-k = 128BLOCK_SIZE_M, BLOCK_SIZE_N, BLOCK_SIZE_K = 128, 256, 64

Kernel source

submission_H100.py82 lines
import triton
import triton.language as tl
import torch

@triton.autotune(
    configs=[
        triton.Config({}, num_warps=8),
    ],
    key=['m', 'n', 'k'],
)
@triton.jit
def matmul_kernel(
    A: torch.Tensor,
    B: torch.Tensor,
    C: torch.Tensor,
    m: int, n: int, k: int,
    BLOCK_SIZE_M: tl.constexpr, BLOCK_SIZE_N: tl.constexpr, BLOCK_SIZE_K: tl.constexpr,
    GROUP_SIZE: tl.constexpr
):
    pid = tl.program_id(0)
    grid_n = tl.cdiv(n, BLOCK_SIZE_N)
    pid_0 = pid // grid_n
    pid_1 = pid % grid_n
    pid_0, pid_1 = tl.swizzle2d(pid_0, pid_1, tl.cdiv(m, BLOCK_SIZE_M), tl.cdiv(n, BLOCK_SIZE_N), GROUP_SIZE)

    a_ptr = tl.make_block_ptr(
        base=A,
        shape=(m, k),
        strides=(k, 1),
        offsets=(pid_0 * BLOCK_SIZE_M, 0),
        block_shape=(BLOCK_SIZE_M, BLOCK_SIZE_K),
        order=(1, 0),
    )

    b_ptr = tl.make_block_ptr(
        base=B,
        shape=(k, n),
        strides=(n, 1),
        offsets=(0, pid_1 * BLOCK_SIZE_N),
        block_shape=(BLOCK_SIZE_K, BLOCK_SIZE_N),
        order=(1, 0),
    )

    accu = tl.zeros((BLOCK_SIZE_M, BLOCK_SIZE_N), dtype=tl.float32)

    for j in range(0, tl.cdiv(k, BLOCK_SIZE_K)):
        a_block = tl.load(a_ptr, boundary_check=(0, 1), padding_option="zero")
        b_block = tl.load(b_ptr, boundary_check=(0, 1), padding_option="zero")
        accu = tl.dot(a_block, b_block, acc=accu)
        a_ptr = tl.advance(a_ptr, (0, BLOCK_SIZE_K))
        b_ptr = tl.advance(b_ptr, (BLOCK_SIZE_K, 0))

    c_ptr = tl.make_block_ptr(
        base=C,
        shape=(m, n),
        strides=(n, 1),
        offsets=(pid_0 * BLOCK_SIZE_M, pid_1 * BLOCK_SIZE_N),
        block_shape=(BLOCK_SIZE_M, BLOCK_SIZE_N),
        order=(1, 0),
    )

    tl.store(c_ptr, accu.to(C.dtype.element_ty), boundary_check=(0, 1))

def custom_kernel(data):

    a, b, c = data

    m, k = a.shape

    _, n = b.shape

    BLOCK_SIZE_M, BLOCK_SIZE_N, BLOCK_SIZE_K = 128, 256, 64

    GROUP_SIZE = 10

    grids = triton.cdiv(m, BLOCK_SIZE_M) * triton.cdiv(n, BLOCK_SIZE_N)

    matmul_kernel[(grids,)](a, b, c, m, n, k, BLOCK_SIZE_M, BLOCK_SIZE_N, BLOCK_SIZE_K, GROUP_SIZE)

    return c

scrolls · 82 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON