Skip to content
KernelIndex
Search⌘K

submission 66458

Joao · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 71 lines, June 9 Researcher Reciprocity License v1.0.

solution5.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectoradd-v2-66458?include=source"
interfacepython
Compatibility
measured onNVIDIA H100
declared hardwareNVIDIA H100
architecturessm_90
dtypesfp16

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
FP16 vector additionsuite of 5 cases
NVIDIA H100
526.0µs
#20 of 44
2025-11-04

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:0876a6f0bb1d3fe537cbb704da2bee23372f416cdef12a0f6863c40ee95a2181
license declaredunknown
license concludedunknown
authorsJoao
imported2026-08-15

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

vector-width = float4__global__ void vectorAdd_float4(const float4* __restrict__ a, const float4* __restrict__ b, float4* __restrict__ c, long long n_float8) {

Kernel source

solution5.py71 lines
#!POPCORN leaderboard vectoradd_v2

import torch
from torch.utils.cpp_extension import load_inline
from typing import List
from task import input_t, output_t
#import sys

add_cuda_source = """

#include <cuda_fp16.h>

__global__ void vectorAdd_float4(const float4*  __restrict__ a, const float4* __restrict__ b, float4* __restrict__ c, long long n_float8) {
    
    // Calculate the global thread ID using a grid-stride loop
    //for (long long i = blockIdx.x * blockDim.x + threadIdx.x; 
    //     i < n_float8; 
    //     i += gridDim.x * blockDim.x) 
    //{
    int i = blockIdx.x * blockDim.x + threadIdx.x; 
        float4 a_vec = a[i];
        float4 b_vec = b[i];
        const half2* a_h = reinterpret_cast<const half2*>(&a_vec);
        const half2* b_h = reinterpret_cast<const half2*>(&b_vec);
        half2 c_h[4];
        c_h[0] = __hadd2(a_h[0], b_h[0]);
        c_h[1] = __hadd2(a_h[1], b_h[1]);
        c_h[2] = __hadd2(a_h[2], b_h[2]);
        c_h[3] = __hadd2(a_h[3], b_h[3]);
        c[i] = *reinterpret_cast<float4*>(c_h);
    //}
}

torch::Tensor add_cuda(torch::Tensor A, torch::Tensor B, torch::Tensor C) {
    int N = A.numel();  
    // N = 268435456
    // n_float8 = N/8 = 33554432
    const long long n_float8 = N / 8;

    const int threads = 256; 
    // blocks = 65536
    const int blocks = (n_float8 + threads - 1) / threads;

    vectorAdd_float4<<<blocks,threads>>>(
        reinterpret_cast<float4*>(A.data_ptr<at::Half>()),
        reinterpret_cast<float4*>(B.data_ptr<at::Half>()),
        reinterpret_cast<float4*>(C.data_ptr<at::Half>()),
        n_float8);

    return C;
}
"""

add_cpp_source = """
#include <torch/extension.h>
#include <cuda_fp16.h>

torch::Tensor add_cuda(torch::Tensor A, torch::Tensor B, torch::Tensor C);
"""

add_module = load_inline(
    name='add_cuda',
    cpp_sources=add_cpp_source,
    cuda_sources=add_cuda_source,
    functions=['add_cuda'],
    verbose=True,
)

def custom_kernel(data: input_t) -> output_t:
    return add_module.add_cuda(data[0], data[1], data[2])
scrolls · 71 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON