submission 776985
ddwinterdd · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 82 lines, June 9 Researcher Reciprocity License v1.0.
submission_H100.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-matmul-v2-776985?include=source"interfacepython
Compatibility
measured onNVIDIA H100
declared hardwareNVIDIA H100
architecturessm_90
dtypesfp16
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:3157a45682b1fe4f844cf12ac4f295932813a070d6390b9a887537100f6d4999
license declaredunknown
license concludedunknown
authorsddwinterdd
imported2026-08-15
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
autotune
@triton.autotune(mma
accu = tl.dot(a_block, b_block, acc=accu)num-warps = 8
triton.Config({}, num_warps=8),tile-k = 128
BLOCK_SIZE_M, BLOCK_SIZE_N, BLOCK_SIZE_K = 128, 256, 64Kernel source
submission_H100.py82 lines
import triton
import triton.language as tl
import torch
@triton.autotune(
configs=[
triton.Config({}, num_warps=8),
],
key=['m', 'n', 'k'],
)
@triton.jit
def matmul_kernel(
A: torch.Tensor,
B: torch.Tensor,
C: torch.Tensor,
m: int, n: int, k: int,
BLOCK_SIZE_M: tl.constexpr, BLOCK_SIZE_N: tl.constexpr, BLOCK_SIZE_K: tl.constexpr,
GROUP_SIZE: tl.constexpr
):
pid = tl.program_id(0)
grid_n = tl.cdiv(n, BLOCK_SIZE_N)
pid_0 = pid // grid_n
pid_1 = pid % grid_n
pid_0, pid_1 = tl.swizzle2d(pid_0, pid_1, tl.cdiv(m, BLOCK_SIZE_M), tl.cdiv(n, BLOCK_SIZE_N), GROUP_SIZE)
a_ptr = tl.make_block_ptr(
base=A,
shape=(m, k),
strides=(k, 1),
offsets=(pid_0 * BLOCK_SIZE_M, 0),
block_shape=(BLOCK_SIZE_M, BLOCK_SIZE_K),
order=(1, 0),
)
b_ptr = tl.make_block_ptr(
base=B,
shape=(k, n),
strides=(n, 1),
offsets=(0, pid_1 * BLOCK_SIZE_N),
block_shape=(BLOCK_SIZE_K, BLOCK_SIZE_N),
order=(1, 0),
)
accu = tl.zeros((BLOCK_SIZE_M, BLOCK_SIZE_N), dtype=tl.float32)
for j in range(0, tl.cdiv(k, BLOCK_SIZE_K)):
a_block = tl.load(a_ptr, boundary_check=(0, 1), padding_option="zero")
b_block = tl.load(b_ptr, boundary_check=(0, 1), padding_option="zero")
accu = tl.dot(a_block, b_block, acc=accu)
a_ptr = tl.advance(a_ptr, (0, BLOCK_SIZE_K))
b_ptr = tl.advance(b_ptr, (BLOCK_SIZE_K, 0))
c_ptr = tl.make_block_ptr(
base=C,
shape=(m, n),
strides=(n, 1),
offsets=(pid_0 * BLOCK_SIZE_M, pid_1 * BLOCK_SIZE_N),
block_shape=(BLOCK_SIZE_M, BLOCK_SIZE_N),
order=(1, 0),
)
tl.store(c_ptr, accu.to(C.dtype.element_ty), boundary_check=(0, 1))
def custom_kernel(data):
a, b, c = data
m, k = a.shape
_, n = b.shape
BLOCK_SIZE_M, BLOCK_SIZE_N, BLOCK_SIZE_K = 128, 256, 64
GROUP_SIZE = 10
grids = triton.cdiv(m, BLOCK_SIZE_M) * triton.cdiv(n, BLOCK_SIZE_N)
matmul_kernel[(grids,)](a, b, c, m, n, k, BLOCK_SIZE_M, BLOCK_SIZE_N, BLOCK_SIZE_K, GROUP_SIZE)
return c
scrolls · 82 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON