submission 777276
ddwinterdd · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 77 lines, June 9 Researcher Reciprocity License v1.0.
submission_A100.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-matmul-v2-777276?include=source"interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp16
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:036ae07a0ccbe4898fe3f59cea202767d2b43694da6167335feba9f2808e5b80
license declaredunknown
license concludedunknown
authorsddwinterdd
imported2026-08-15
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
mma
accu = tl.dot(a_block, b_block, acc=accu)tile-k = 128
BLOCK_SIZE_M, BLOCK_SIZE_N, BLOCK_SIZE_K = 128, 128, 32Kernel source
submission_A100.py77 lines
import triton
import triton.language as tl
import torch
@triton.jit
def matmul_kernel(
A: torch.Tensor,
B: torch.Tensor,
C: torch.Tensor,
m: int, n: int, k: int,
BLOCK_SIZE_M: tl.constexpr, BLOCK_SIZE_N: tl.constexpr, BLOCK_SIZE_K: tl.constexpr,
GROUP_SIZE: tl.constexpr
):
pid = tl.program_id(0)
grid_n = tl.cdiv(n, BLOCK_SIZE_N)
pid_0 = pid // grid_n
pid_1 = pid % grid_n
pid_0, pid_1 = tl.swizzle2d(pid_0, pid_1, tl.cdiv(m, BLOCK_SIZE_M), tl.cdiv(n, BLOCK_SIZE_N), GROUP_SIZE)
a_ptr = tl.make_block_ptr(
base=A,
shape=(m, k),
strides=(k, 1),
offsets=(pid_0 * BLOCK_SIZE_M, 0),
block_shape=(BLOCK_SIZE_M, BLOCK_SIZE_K),
order=(1, 0),
)
b_ptr = tl.make_block_ptr(
base=B,
shape=(k, n),
strides=(n, 1),
offsets=(0, pid_1 * BLOCK_SIZE_N),
block_shape=(BLOCK_SIZE_K, BLOCK_SIZE_N),
order=(1, 0),
)
accu = tl.zeros((BLOCK_SIZE_M, BLOCK_SIZE_N), dtype=tl.float32)
for j in range(0, tl.cdiv(k, BLOCK_SIZE_K)):
a_block = tl.load(a_ptr, boundary_check=(0, 1), padding_option="zero")
b_block = tl.load(b_ptr, boundary_check=(0, 1), padding_option="zero")
accu = tl.dot(a_block, b_block, acc=accu)
a_ptr = tl.advance(a_ptr, (0, BLOCK_SIZE_K))
b_ptr = tl.advance(b_ptr, (BLOCK_SIZE_K, 0))
c_ptr = tl.make_block_ptr(
base=C,
shape=(m, n),
strides=(n, 1),
offsets=(pid_0 * BLOCK_SIZE_M, pid_1 * BLOCK_SIZE_N),
block_shape=(BLOCK_SIZE_M, BLOCK_SIZE_N),
order=(1, 0),
)
tl.store(c_ptr, accu.to(C.dtype.element_ty), boundary_check=(0, 1))
def custom_kernel(data):
a, b, c = data
m, k = a.shape
_, n = b.shape
BLOCK_SIZE_M, BLOCK_SIZE_N, BLOCK_SIZE_K = 128, 128, 32
GROUP_SIZE = 12
grids = triton.cdiv(m, BLOCK_SIZE_M) * triton.cdiv(n, BLOCK_SIZE_N)
matmul_kernel[(grids,)](a, b, c, m, n, k, BLOCK_SIZE_M, BLOCK_SIZE_N, BLOCK_SIZE_K, GROUP_SIZE)
return c
scrolls · 77 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Changes from previous submission
Against this author's previous submission submission 776985.
⋯ 1 unchanged linesimport triton.language as tlimport torch- @triton.autotune(- configs=[- triton.Config({}, num_warps=8),- ],- key=['m', 'n', 'k'],- )+@triton.jitdef matmul_kernel(A: torch.Tensor,⋯ 55 unchanged lines_, n = b.shape- BLOCK_SIZE_M, BLOCK_SIZE_N, BLOCK_SIZE_K = 128, 256, 64+ BLOCK_SIZE_M, BLOCK_SIZE_N, BLOCK_SIZE_K = 128, 128, 32- GROUP_SIZE = 10+ GROUP_SIZE = 12grids = triton.cdiv(m, BLOCK_SIZE_M) * triton.cdiv(n, BLOCK_SIZE_N)
scrolls · 26 diff lines total
Best evidence level for this revision: reported
JSON