Skip to content
KernelIndex
Search⌘K

submission 596202

kirpar · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 88 lines, June 9 Researcher Reciprocity License v1.0.

submission_my.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectorsum-v2-596202?include=source"
interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
Vector sum reductionsuite of 6 cases
NVIDIA A100
158.8µs
#65 of 96
2026-03-20

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:08ef98d7ea84eafd72fd9e18733203b4cae01fb991aa058f893cb159f1b894e1
license declaredunknown
license concludedunknown
authorskirpar
imported2026-08-15

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

shared-memoryextern __shared__ float sdata[];
vector-width = float4const float4* input_float4 = reinterpret_cast<const float4*>(input);

Kernel source

submission_my.py88 lines
#!POPCORN leaderboard vectorsum_v2
#!POPCORN gpu A100

import os
os.environ["TORCH_EXTENSIONS_DIR"] = "/tmp/torch_extensions" 

import torch
from torch.utils.cpp_extension import load_inline
from task import input_t, output_t

cuda_source = """
#include <cuda_runtime.h>

__global__ void vectorSumKernel(const float* __restrict__ input, float* __restrict__ output, int N) {
    extern __shared__ float sdata[];
    unsigned int local_id = threadIdx.x;
    unsigned int global_id = blockIdx.x * blockDim.x + threadIdx.x;
    unsigned int grid_stride = blockDim.x * gridDim.x;

    float thread_local_sum = 0.0f;
    int N_vec = N / 4; 
    const float4* input_float4 = reinterpret_cast<const float4*>(input);

    for (int i = global_id; i < N_vec; i += grid_stride) {
        float4 vec = input_float4[i];
        thread_local_sum += vec.x + vec.y + vec.z + vec.w;
    }

    int tail_start = N_vec * 4;
    for (int i = tail_start + global_id; i < N; i += grid_stride) {
        thread_local_sum += input[i];
    }

    sdata[local_id] = thread_local_sum;
    __syncthreads(); 

    for (unsigned int stride = blockDim.x / 2; stride > 0; stride >>= 1) {
        if (local_id < stride) {
            sdata[local_id] += sdata[local_id + stride];
        }
        __syncthreads(); 
    }

    if (local_id == 0) {
        atomicAdd(output, sdata[0]);
    }
}

// NOTE: We changed the wrapper to accept the pre-allocated output tensor
void vector_sum(torch::Tensor input, torch::Tensor output) {
    int N = input.numel();
    int threads = 256;
    int blocks = std::min((N + threads - 1) / threads, 1024); 
    int shared_mem_bytes = threads * sizeof(float);
    
    // We must zero out their output tensor before using atomicAdd!
    output.zero_();
    
    vectorSumKernel<<<blocks, threads, shared_mem_bytes>>>(
        input.data_ptr<float>(), 
        output.data_ptr<float>(), 
        N
    );
}
"""

cpp_source = "void vector_sum(torch::Tensor input, torch::Tensor output);"

vector_sum_module = load_inline(
    name="custom_vector_sum",
    cpp_sources=cpp_source,
    cuda_sources=cuda_source,
    functions=["vector_sum"],
    extra_cflags=["-O3"],
    extra_cuda_cflags=["-O3", "--use_fast_math"]
)

# We disable Dynamo here too just to be safe
@torch.compiler.disable
def custom_kernel(data: input_t) -> output_t:
    input_tensor, output_tensor = data
    
    # Run our C++ function using their pre-allocated tensors
    vector_sum_module.vector_sum(input_tensor, output_tensor)
    
    return output_tensor[0]

    
scrolls · 88 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON