Skip to content
KernelIndex
Search⌘K

submission 66711

P · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 82 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectorsum-v2-66711?include=source"
interfacepython
Compatibility
measured onNVIDIA L4
declared hardwareNVIDIA L4
architecturessm_89
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
Vector sum reductionsuite of 6 cases
NVIDIA L4
1.66ms
#24 of 26
2025-11-05

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:347d1dfe2705623e29c8c380f0a897329353236a7a511466af103ba254ebfbeb
license declaredunknown
license concludedunknown
authorsP
imported2026-08-15

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

shared-memory__shared__ scalar_t partial_sum[BLOCK_SIZE];

Kernel source

submission.py82 lines
import torch
from utils import DeterministicContext
from torch.utils.cpp_extension import load_inline
from typing import List
from task import input_t, output_t

sum_cuda_source = """
#define BLOCK_SIZE 1024
template <typename scalar_t>
__global__ void sum_kernel(const scalar_t* __restrict__ A,
                           scalar_t* __restrict__ B,
                           int N) {

    __shared__ scalar_t partial_sum[BLOCK_SIZE];

    unsigned int tid = threadIdx.x;
    unsigned int i = blockIdx.x * (BLOCK_SIZE) + tid;

    if (i < N) {
        partial_sum[tid] = A[i];
    }
    else {
        partial_sum[tid] = 0;
    }

    for (unsigned int stride = BLOCK_SIZE/2; stride >= 1; stride /= 2) {
        __syncthreads();
        if (tid < stride) {
            partial_sum[tid] += partial_sum[tid + stride];
        }
    }
    __syncthreads();

    if (tid == 0) {
        atomicAdd(B, partial_sum[0]);
    }
}

torch::Tensor sum_cuda(torch::Tensor A, torch::Tensor B) {
    int N = A.numel();

    int blocks = (N + BLOCK_SIZE - 1) / BLOCK_SIZE;

    sum_kernel<float><<<blocks, BLOCK_SIZE>>>(
        A.data_ptr<float>(),
        B.data_ptr<float>(),
        N
    );

    return B;
}
"""

sum_module = load_inline(
    name='sum_cuda_ext',
    cpp_sources="torch::Tensor sum_cuda(torch::Tensor A, torch::Tensor B);",
    cuda_sources=sum_cuda_source,
    functions=['sum_cuda'],
    verbose=True,
)

def sum(A, B):
    if not A.is_cuda or not B.is_cuda:
        raise RuntimeError("Three tensors must be on GPU")
    return sum_module.sum_cuda(A, B)

def custom_kernel(data: input_t) -> output_t:
    """
    Custom implementation of vector addition using CUDA.
    Args:
        inputs: List of pairs of tensors [A, B] to be added.
    Returns:
        Tensor containing element-wise sum.
    """
    A, B = data
    B.zero_()
    assert A.is_cuda and B.is_cuda, "Input tensors must be on GPU"

    # Simply reuse the existing add function we already defined
    # This avoids the compilation issues with the inline kernel
    return sum(A, B)[0]
scrolls · 82 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Changes from previous submission

Against this author's previous submission submission 66709.

Best evidence level for this revision: reported

JSON