submission 66711
P · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 82 lines, June 9 Researcher Reciprocity License v1.0.
submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectorsum-v2-66711?include=source"interfacepython
Compatibility
measured onNVIDIA L4
declared hardwareNVIDIA L4
architecturessm_89
dtypesfp32
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:347d1dfe2705623e29c8c380f0a897329353236a7a511466af103ba254ebfbeb
license declaredunknown
license concludedunknown
authorsP
imported2026-08-15
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
shared-memory
__shared__ scalar_t partial_sum[BLOCK_SIZE];Kernel source
submission.py82 lines
import torch
from utils import DeterministicContext
from torch.utils.cpp_extension import load_inline
from typing import List
from task import input_t, output_t
sum_cuda_source = """
#define BLOCK_SIZE 1024
template <typename scalar_t>
__global__ void sum_kernel(const scalar_t* __restrict__ A,
scalar_t* __restrict__ B,
int N) {
__shared__ scalar_t partial_sum[BLOCK_SIZE];
unsigned int tid = threadIdx.x;
unsigned int i = blockIdx.x * (BLOCK_SIZE) + tid;
if (i < N) {
partial_sum[tid] = A[i];
}
else {
partial_sum[tid] = 0;
}
for (unsigned int stride = BLOCK_SIZE/2; stride >= 1; stride /= 2) {
__syncthreads();
if (tid < stride) {
partial_sum[tid] += partial_sum[tid + stride];
}
}
__syncthreads();
if (tid == 0) {
atomicAdd(B, partial_sum[0]);
}
}
torch::Tensor sum_cuda(torch::Tensor A, torch::Tensor B) {
int N = A.numel();
int blocks = (N + BLOCK_SIZE - 1) / BLOCK_SIZE;
sum_kernel<float><<<blocks, BLOCK_SIZE>>>(
A.data_ptr<float>(),
B.data_ptr<float>(),
N
);
return B;
}
"""
sum_module = load_inline(
name='sum_cuda_ext',
cpp_sources="torch::Tensor sum_cuda(torch::Tensor A, torch::Tensor B);",
cuda_sources=sum_cuda_source,
functions=['sum_cuda'],
verbose=True,
)
def sum(A, B):
if not A.is_cuda or not B.is_cuda:
raise RuntimeError("Three tensors must be on GPU")
return sum_module.sum_cuda(A, B)
def custom_kernel(data: input_t) -> output_t:
"""
Custom implementation of vector addition using CUDA.
Args:
inputs: List of pairs of tensors [A, B] to be added.
Returns:
Tensor containing element-wise sum.
"""
A, B = data
B.zero_()
assert A.is_cuda and B.is_cuda, "Input tensors must be on GPU"
# Simply reuse the existing add function we already defined
# This avoids the compilation issues with the inline kernel
return sum(A, B)[0]
scrolls · 82 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Changes from previous submission
Against this author's previous submission submission 66709.
Best evidence level for this revision: reported
JSON