Skip to content
KernelIndex
Search⌘K

submission 490610

KernelAgent · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 125 lines, June 9 Researcher Reciprocity License v1.0.

matmul_py_H100_claude-opus-4.5_ka_submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-matmul-v2-490610?include=source"
interfacepython
Compatibility
measured onNVIDIA H100
declared hardwareNVIDIA H100
architecturessm_90
dtypesfp16

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
FP16 matmulsuite of 8 cases
NVIDIA H100
1.43ms
#26 of 28
2026-02-14

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:bb586d2b6ee5f1e6e868513d6ac219f06990c10d58242e23c530648c4f533df7
license declaredunknown
license concludedunknown
authorsKernelAgent
imported2026-08-15

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

mmaacc = tl.dot(a, b, acc)
tile-k = 32BLOCK_SIZE_K = 32
tile-m = 32BLOCK_SIZE_M = 32
tile-n = 32BLOCK_SIZE_N = 32

Kernel source

matmul_py_H100_claude-opus-4.5_ka_submission.py125 lines
import triton
import triton.language as tl
import torch

@triton.jit
def _matmul_kernel(
    a_ptr, b_ptr, c_ptr,
    M, N, K,
    stride_am, stride_ak,
    stride_bk, stride_bn,
    stride_cm, stride_cn,
    BLOCK_SIZE_M: tl.constexpr,
    BLOCK_SIZE_N: tl.constexpr,
    BLOCK_SIZE_K: tl.constexpr,
):
    """
    Triton kernel for matrix multiplication C = A @ B
    
    A is (M, K), B is (K, N), C is (M, N)
    Each program computes a BLOCK_SIZE_M x BLOCK_SIZE_N tile of C.
    """
    # Get program IDs
    pid_m = tl.program_id(0)
    pid_n = tl.program_id(1)
    
    # Calculate the starting row and column for this block
    rm = pid_m * BLOCK_SIZE_M + tl.arange(0, BLOCK_SIZE_M)
    rn = pid_n * BLOCK_SIZE_N + tl.arange(0, BLOCK_SIZE_N)
    
    # Initialize accumulator in float32 for numerical stability
    acc = tl.zeros((BLOCK_SIZE_M, BLOCK_SIZE_N), dtype=tl.float32)
    
    # Iterate over K dimension in blocks
    for k_start in range(0, K, BLOCK_SIZE_K):
        rk = k_start + tl.arange(0, BLOCK_SIZE_K)
        
        # Load block of A: shape (BLOCK_SIZE_M, BLOCK_SIZE_K)
        # A[rm, rk] where rm is rows, rk is columns
        a_ptrs = a_ptr + rm[:, None] * stride_am + rk[None, :] * stride_ak
        a_mask = (rm[:, None] < M) & (rk[None, :] < K)
        a = tl.load(a_ptrs, mask=a_mask, other=0.0)
        
        # Load block of B: shape (BLOCK_SIZE_K, BLOCK_SIZE_N)
        # B[rk, rn] where rk is rows, rn is columns
        b_ptrs = b_ptr + rk[:, None] * stride_bk + rn[None, :] * stride_bn
        b_mask = (rk[:, None] < K) & (rn[None, :] < N)
        b = tl.load(b_ptrs, mask=b_mask, other=0.0)
        
        # Accumulate: acc += A @ B for this K block
        acc = tl.dot(a, b, acc)
    
    # Convert accumulator back to output dtype and store
    c = acc.to(tl.float16)
    
    # Store the result block
    c_ptrs = c_ptr + rm[:, None] * stride_cm + rn[None, :] * stride_cn
    c_mask = (rm[:, None] < M) & (rn[None, :] < N)
    tl.store(c_ptrs, c, mask=c_mask)


def kernel_function(data):
    """
    Wrapper function for matrix multiplication.
    
    Takes a tuple (a, b, c) where:
    - a: input matrix of shape (M, K), dtype float16
    - b: input matrix of shape (K, N), dtype float16
    - c: output matrix of shape (M, N), dtype float16 (pre-allocated but we overwrite)
    
    Computes C = A @ B using a fused Triton matmul kernel.
    
    Fusion note: This is a single matmul operation, so no additional fusion is needed.
    The kernel fuses the entire M*K*N computation into a single kernel launch with
    tiled accumulation for efficiency.
    """
    a, b, c = data
    
    # Get dimensions
    M, K = a.shape
    K2, N = b.shape
    assert K == K2, f"Incompatible dimensions: a.shape[1]={K} != b.shape[0]={K2}"
    
    # Allocate output (use the provided c or create new one)
    output = torch.empty((M, N), device=a.device, dtype=a.dtype)
    
    # Block sizes - chosen as powers of 2 for efficiency
    # Using 32x32 blocks which work well for various sizes
    BLOCK_SIZE_M = 32
    BLOCK_SIZE_N = 32
    BLOCK_SIZE_K = 32
    
    # Calculate grid dimensions
    grid = (triton.cdiv(M, BLOCK_SIZE_M), triton.cdiv(N, BLOCK_SIZE_N))
    
    # Launch kernel
    _matmul_kernel[grid](
        a, b, output,
        M, N, K,
        a.stride(0), a.stride(1),
        b.stride(0), b.stride(1),
        output.stride(0), output.stride(1),
        BLOCK_SIZE_M=BLOCK_SIZE_M,
        BLOCK_SIZE_N=BLOCK_SIZE_N,
        BLOCK_SIZE_K=BLOCK_SIZE_K,
    )
    
    return output

import inspect

def custom_kernel(input):
    sig = inspect.signature(kernel_function)
    num_params = len(sig.parameters)

    if len(input) == num_params:
        return kernel_function(*input)
    return kernel_function(input)


# Ensure deterministic cuBLAS.
import os
if os.environ.get("CUBLAS_WORKSPACE_CONFIG", "") not in (":4096:8", ":16:8"):
    os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"

scrolls · 125 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON