Skip to content
KernelIndex
Search⌘K

submission 116763

achal · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 51 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectoradd-v2-116763?include=source"
interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp16

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
FP16 vector additionsuite of 5 cases
NVIDIA A100
75.5ms
#72 of 87
2025-11-30

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:4d33fc944f4402da17375fdebcc2eb1fe1b8b4c33751eeb33cf44a3d81525357
license declaredunknown
license concludedunknown
authorsachal
imported2026-08-15

Kernel source

submission.py51 lines
from task import input_t, output_t

import cutlass
import cutlass.cute as cute
from cutlass.cute.runtime import make_ptr
import torch

@cute.kernel
def elementwise_add_kernel(M, N, gX: cute.Tensor, gY: cute.Tensor, gZ: cute.Tensor):
    tidx, tidy, _ = cute.arch.thread_idx()
    bidx, bidy, _ = cute.arch.block_idx()
    bdimx, bdimy, _ = cute.arch.block_dim()

    tx = bdimx*bidx + tidx
    ty = bdimy*bidy + tidy
    if tx < N and ty < M:
        gZ[ty, tx] = gX[ty, tx] + gY[ty, tx]

@cute.jit
def elementwise_add(M, N, X: cute.Tensor, Y: cute.Tensor, Z: cute.Tensor):
    block_dim = (32, 32, 1)
    grid_dim = (
        (N + block_dim[0] - 1)//block_dim[0],
        (M + block_dim[1] - 1)//block_dim[1],
        1,
    )
    elementwise_add_kernel(M, N, X, Y, Z).launch(grid=grid_dim, block=block_dim)

_compiled_function_cache = None

def compile_function(fn, A, B, C):
    compiled_function_cache = globals()["_compiled_function_cache"]
    if compiled_function_cache is not None:
        return compiled_function_cache
    
    compiled_function_cache = cute.compile(fn, A.shape[0], A.shape[1], A, B, C)
    return compiled_function_cache

def custom_kernel_(data: input_t) -> output_t:
    A, B, C = data
    
    A_ = cute.runtime.from_dlpack(A)
    B_ = cute.runtime.from_dlpack(B)
    C_ = cute.runtime.from_dlpack(C)
    compiled_function = compile_function(elementwise_add, A_, B_, C_)
    
    compiled_function(A.shape[0], A.shape[1], A_, B_, C_)
    
    return C

custom_kernel = custom_kernel_
scrolls · 51 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON