submission 490610
KernelAgent · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 125 lines, June 9 Researcher Reciprocity License v1.0.
matmul_py_H100_claude-opus-4.5_ka_submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-matmul-v2-490610?include=source"interfacepython
Compatibility
measured onNVIDIA H100
declared hardwareNVIDIA H100
architecturessm_90
dtypesfp16
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:bb586d2b6ee5f1e6e868513d6ac219f06990c10d58242e23c530648c4f533df7
license declaredunknown
license concludedunknown
authorsKernelAgent
imported2026-08-15
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
mma
acc = tl.dot(a, b, acc)tile-k = 32
BLOCK_SIZE_K = 32tile-m = 32
BLOCK_SIZE_M = 32tile-n = 32
BLOCK_SIZE_N = 32Kernel source
matmul_py_H100_claude-opus-4.5_ka_submission.py125 lines
import triton
import triton.language as tl
import torch
@triton.jit
def _matmul_kernel(
a_ptr, b_ptr, c_ptr,
M, N, K,
stride_am, stride_ak,
stride_bk, stride_bn,
stride_cm, stride_cn,
BLOCK_SIZE_M: tl.constexpr,
BLOCK_SIZE_N: tl.constexpr,
BLOCK_SIZE_K: tl.constexpr,
):
"""
Triton kernel for matrix multiplication C = A @ B
A is (M, K), B is (K, N), C is (M, N)
Each program computes a BLOCK_SIZE_M x BLOCK_SIZE_N tile of C.
"""
# Get program IDs
pid_m = tl.program_id(0)
pid_n = tl.program_id(1)
# Calculate the starting row and column for this block
rm = pid_m * BLOCK_SIZE_M + tl.arange(0, BLOCK_SIZE_M)
rn = pid_n * BLOCK_SIZE_N + tl.arange(0, BLOCK_SIZE_N)
# Initialize accumulator in float32 for numerical stability
acc = tl.zeros((BLOCK_SIZE_M, BLOCK_SIZE_N), dtype=tl.float32)
# Iterate over K dimension in blocks
for k_start in range(0, K, BLOCK_SIZE_K):
rk = k_start + tl.arange(0, BLOCK_SIZE_K)
# Load block of A: shape (BLOCK_SIZE_M, BLOCK_SIZE_K)
# A[rm, rk] where rm is rows, rk is columns
a_ptrs = a_ptr + rm[:, None] * stride_am + rk[None, :] * stride_ak
a_mask = (rm[:, None] < M) & (rk[None, :] < K)
a = tl.load(a_ptrs, mask=a_mask, other=0.0)
# Load block of B: shape (BLOCK_SIZE_K, BLOCK_SIZE_N)
# B[rk, rn] where rk is rows, rn is columns
b_ptrs = b_ptr + rk[:, None] * stride_bk + rn[None, :] * stride_bn
b_mask = (rk[:, None] < K) & (rn[None, :] < N)
b = tl.load(b_ptrs, mask=b_mask, other=0.0)
# Accumulate: acc += A @ B for this K block
acc = tl.dot(a, b, acc)
# Convert accumulator back to output dtype and store
c = acc.to(tl.float16)
# Store the result block
c_ptrs = c_ptr + rm[:, None] * stride_cm + rn[None, :] * stride_cn
c_mask = (rm[:, None] < M) & (rn[None, :] < N)
tl.store(c_ptrs, c, mask=c_mask)
def kernel_function(data):
"""
Wrapper function for matrix multiplication.
Takes a tuple (a, b, c) where:
- a: input matrix of shape (M, K), dtype float16
- b: input matrix of shape (K, N), dtype float16
- c: output matrix of shape (M, N), dtype float16 (pre-allocated but we overwrite)
Computes C = A @ B using a fused Triton matmul kernel.
Fusion note: This is a single matmul operation, so no additional fusion is needed.
The kernel fuses the entire M*K*N computation into a single kernel launch with
tiled accumulation for efficiency.
"""
a, b, c = data
# Get dimensions
M, K = a.shape
K2, N = b.shape
assert K == K2, f"Incompatible dimensions: a.shape[1]={K} != b.shape[0]={K2}"
# Allocate output (use the provided c or create new one)
output = torch.empty((M, N), device=a.device, dtype=a.dtype)
# Block sizes - chosen as powers of 2 for efficiency
# Using 32x32 blocks which work well for various sizes
BLOCK_SIZE_M = 32
BLOCK_SIZE_N = 32
BLOCK_SIZE_K = 32
# Calculate grid dimensions
grid = (triton.cdiv(M, BLOCK_SIZE_M), triton.cdiv(N, BLOCK_SIZE_N))
# Launch kernel
_matmul_kernel[grid](
a, b, output,
M, N, K,
a.stride(0), a.stride(1),
b.stride(0), b.stride(1),
output.stride(0), output.stride(1),
BLOCK_SIZE_M=BLOCK_SIZE_M,
BLOCK_SIZE_N=BLOCK_SIZE_N,
BLOCK_SIZE_K=BLOCK_SIZE_K,
)
return output
import inspect
def custom_kernel(input):
sig = inspect.signature(kernel_function)
num_params = len(sig.parameters)
if len(input) == num_params:
return kernel_function(*input)
return kernel_function(input)
# Ensure deterministic cuBLAS.
import os
if os.environ.get("CUBLAS_WORKSPACE_CONFIG", "") not in (":4096:8", ":16:8"):
os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"
scrolls · 125 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON