submission 768762
sophia · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 29 lines, June 9 Researcher Reciprocity License v1.0.
aaaaf5c09e19.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-prefixsum-v2-768762?include=source"interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp32
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:3d33bf2f85f7eb4bf9600a71361693367e7d8cbd1f071425a78044dd9bcb6519
license declaredunknown
license concludedunknown
authorssophia
imported2026-08-15
Kernel source
aaaaf5c09e19.py29 lines
#!POPCORN leaderboard prefixsum_v2
import torch
from task import input_t, output_t
CUDA_SOURCE = r"""
// naive serial prefix sum — one thread does all the work
__global__ void solution(const float* __restrict__ input, float* __restrict__ output, int n) {
if (threadIdx.x == 0 && blockIdx.x == 0) {
double acc = 0.0;
for (int i = 0; i < n; i++) {
acc += (double)input[i];
output[i] = (float)acc;
}
}
}
"""
_kernel = torch.cuda._compile_kernel(CUDA_SOURCE, "solution")
def custom_kernel(data: input_t) -> output_t:
x, y = data
n = x.numel()
# adjust grid/block to match your kernel
threads = (256, 1, 1)
blocks = ((n + 255) // 256, 1, 1)
_kernel(blocks, threads, (x, y, n))
return y
scrolls · 29 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON