Skip to content
KernelIndex
Search⌘K

submission 777276

ddwinterdd · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 77 lines, June 9 Researcher Reciprocity License v1.0.

submission_A100.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-matmul-v2-777276?include=source"
interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp16

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
FP16 matmulsuite of 8 cases
NVIDIA A100
825.7µs
#25 of 27
2026-04-20

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:036ae07a0ccbe4898fe3f59cea202767d2b43694da6167335feba9f2808e5b80
license declaredunknown
license concludedunknown
authorsddwinterdd
imported2026-08-15

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

mmaaccu = tl.dot(a_block, b_block, acc=accu)
tile-k = 128BLOCK_SIZE_M, BLOCK_SIZE_N, BLOCK_SIZE_K = 128, 128, 32

Kernel source

submission_A100.py77 lines
import triton
import triton.language as tl
import torch


@triton.jit
def matmul_kernel(
    A: torch.Tensor,
    B: torch.Tensor,
    C: torch.Tensor,
    m: int, n: int, k: int,
    BLOCK_SIZE_M: tl.constexpr, BLOCK_SIZE_N: tl.constexpr, BLOCK_SIZE_K: tl.constexpr,
    GROUP_SIZE: tl.constexpr
):
    pid = tl.program_id(0)
    grid_n = tl.cdiv(n, BLOCK_SIZE_N)
    pid_0 = pid // grid_n
    pid_1 = pid % grid_n
    pid_0, pid_1 = tl.swizzle2d(pid_0, pid_1, tl.cdiv(m, BLOCK_SIZE_M), tl.cdiv(n, BLOCK_SIZE_N), GROUP_SIZE)

    a_ptr = tl.make_block_ptr(
        base=A,
        shape=(m, k),
        strides=(k, 1),
        offsets=(pid_0 * BLOCK_SIZE_M, 0),
        block_shape=(BLOCK_SIZE_M, BLOCK_SIZE_K),
        order=(1, 0),
    )

    b_ptr = tl.make_block_ptr(
        base=B,
        shape=(k, n),
        strides=(n, 1),
        offsets=(0, pid_1 * BLOCK_SIZE_N),
        block_shape=(BLOCK_SIZE_K, BLOCK_SIZE_N),
        order=(1, 0),
    )

    accu = tl.zeros((BLOCK_SIZE_M, BLOCK_SIZE_N), dtype=tl.float32)

    for j in range(0, tl.cdiv(k, BLOCK_SIZE_K)):
        a_block = tl.load(a_ptr, boundary_check=(0, 1), padding_option="zero")
        b_block = tl.load(b_ptr, boundary_check=(0, 1), padding_option="zero")
        accu = tl.dot(a_block, b_block, acc=accu)
        a_ptr = tl.advance(a_ptr, (0, BLOCK_SIZE_K))
        b_ptr = tl.advance(b_ptr, (BLOCK_SIZE_K, 0))

    c_ptr = tl.make_block_ptr(
        base=C,
        shape=(m, n),
        strides=(n, 1),
        offsets=(pid_0 * BLOCK_SIZE_M, pid_1 * BLOCK_SIZE_N),
        block_shape=(BLOCK_SIZE_M, BLOCK_SIZE_N),
        order=(1, 0),
    )

    tl.store(c_ptr, accu.to(C.dtype.element_ty), boundary_check=(0, 1))

def custom_kernel(data):

    a, b, c = data

    m, k = a.shape

    _, n = b.shape

    BLOCK_SIZE_M, BLOCK_SIZE_N, BLOCK_SIZE_K = 128, 128, 32

    GROUP_SIZE = 12

    grids = triton.cdiv(m, BLOCK_SIZE_M) * triton.cdiv(n, BLOCK_SIZE_N)

    matmul_kernel[(grids,)](a, b, c, m, n, k, BLOCK_SIZE_M, BLOCK_SIZE_N, BLOCK_SIZE_K, GROUP_SIZE)

    return c

scrolls · 77 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Changes from previous submission

Against this author's previous submission submission 776985.

⋯ 1 unchanged lines
import triton.language as tl
import torch
- @triton.autotune(
- configs=[
- triton.Config({}, num_warps=8),
- ],
- key=['m', 'n', 'k'],
- )
+
@triton.jit
def matmul_kernel(
A: torch.Tensor,
⋯ 55 unchanged lines
_, n = b.shape
- BLOCK_SIZE_M, BLOCK_SIZE_N, BLOCK_SIZE_K = 128, 256, 64
+ BLOCK_SIZE_M, BLOCK_SIZE_N, BLOCK_SIZE_K = 128, 128, 32
- GROUP_SIZE = 10
+ GROUP_SIZE = 12
grids = triton.cdiv(m, BLOCK_SIZE_M) * triton.cdiv(n, BLOCK_SIZE_N)
scrolls · 26 diff lines total

Best evidence level for this revision: reported

JSON