submission 772562
ddwinterdd · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 59 lines, June 9 Researcher Reciprocity License v1.0.
submission_1_3.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectoradd-v2-772562?include=source"interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp16
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:8a67247a3d41009056dcdb183cbc60a67f6853732e620e37293a96f35867e6a7
license declaredunknown
license concludedunknown
authorsddwinterdd
imported2026-08-15
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
autotune
@triton.autotune(num-warps = 16
triton.Config({}, num_warps=16),Kernel source
submission_1_3.py59 lines
import triton
import triton.language as tl
import torch
from typing import Tuple
# Reference input/output type aliases (same as in reference.py)
# input_t = Tuple[torch.Tensor, torch.Tensor, torch.Tensor]
# output_t = torch.Tensor
@triton.autotune(
configs=[
triton.Config({}, num_warps=16),
],
key=['N'],
)
@triton.jit
def add_kernel(A, B, C, N, BLOCK_SIZE: tl.constexpr):
"""Element‑wise addition of two 1‑D flattened tensors.
A, B, C are pointers to the start of the tensors in device memory.
N is the total number of elements.
BLOCK_SIZE is the number of elements processed per program.
"""
pid = tl.program_id(0)
offs = pid * BLOCK_SIZE + tl.arange(0, BLOCK_SIZE)
mask = offs < N
a = tl.load(A + offs, mask=mask)
b = tl.load(B + offs, mask=mask)
c = a + b
tl.store(C + offs, c, mask=mask)
def custom_kernel(data: Tuple[torch.Tensor, torch.Tensor, torch.Tensor]) -> torch.Tensor:
"""Entry point matching the reference signature.
Args:
data: (A, B, output) where A, B, output are torch.float16 CUDA tensors of shape (N, N).
Returns:
The output tensor containing A + B.
"""
A, B, out = data
assert A.is_cuda and B.is_cuda and out.is_cuda, "All tensors must be on CUDA"
assert A.dtype == torch.float16 and B.dtype == torch.float16 and out.dtype == torch.float16, "All tensors must be float16"
assert A.shape == B.shape == out.shape, "All tensors must have the same shape"
# Flatten tensors for a 1‑D launch
A_flat = A.view(-1)
B_flat = B.view(-1)
out_flat = out.view(-1)
total_elements = A_flat.numel()
# Choose a block size that is a power of two and fits into shared memory
BLOCK_SIZE = 4096
grid = lambda meta: ( (total_elements + meta['BLOCK_SIZE'] - 1) // meta['BLOCK_SIZE'], )
add_kernel[grid](A_flat, B_flat, out_flat, total_elements, BLOCK_SIZE=BLOCK_SIZE)
return out
scrolls · 59 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON