submission 116763
achal · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 51 lines, June 9 Researcher Reciprocity License v1.0.
submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectoradd-v2-116763?include=source"interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp16
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:4d33fc944f4402da17375fdebcc2eb1fe1b8b4c33751eeb33cf44a3d81525357
license declaredunknown
license concludedunknown
authorsachal
imported2026-08-15
Kernel source
submission.py51 lines
from task import input_t, output_t
import cutlass
import cutlass.cute as cute
from cutlass.cute.runtime import make_ptr
import torch
@cute.kernel
def elementwise_add_kernel(M, N, gX: cute.Tensor, gY: cute.Tensor, gZ: cute.Tensor):
tidx, tidy, _ = cute.arch.thread_idx()
bidx, bidy, _ = cute.arch.block_idx()
bdimx, bdimy, _ = cute.arch.block_dim()
tx = bdimx*bidx + tidx
ty = bdimy*bidy + tidy
if tx < N and ty < M:
gZ[ty, tx] = gX[ty, tx] + gY[ty, tx]
@cute.jit
def elementwise_add(M, N, X: cute.Tensor, Y: cute.Tensor, Z: cute.Tensor):
block_dim = (32, 32, 1)
grid_dim = (
(N + block_dim[0] - 1)//block_dim[0],
(M + block_dim[1] - 1)//block_dim[1],
1,
)
elementwise_add_kernel(M, N, X, Y, Z).launch(grid=grid_dim, block=block_dim)
_compiled_function_cache = None
def compile_function(fn, A, B, C):
compiled_function_cache = globals()["_compiled_function_cache"]
if compiled_function_cache is not None:
return compiled_function_cache
compiled_function_cache = cute.compile(fn, A.shape[0], A.shape[1], A, B, C)
return compiled_function_cache
def custom_kernel_(data: input_t) -> output_t:
A, B, C = data
A_ = cute.runtime.from_dlpack(A)
B_ = cute.runtime.from_dlpack(B)
C_ = cute.runtime.from_dlpack(C)
compiled_function = compile_function(elementwise_add, A_, B_, C_)
compiled_function(A.shape[0], A.shape[1], A_, B_, C_)
return C
custom_kernel = custom_kernel_scrolls · 51 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON