Skip to content
KernelIndex
Search⌘K

submission 768762

sophia · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 29 lines, June 9 Researcher Reciprocity License v1.0.

aaaaf5c09e19.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-prefixsum-v2-768762?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
Inclusive prefix sumsuite of 11 cases
NVIDIA B200
8.85s
#23 of 23
2026-04-15

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:3d33bf2f85f7eb4bf9600a71361693367e7d8cbd1f071425a78044dd9bcb6519
license declaredunknown
license concludedunknown
authorssophia
imported2026-08-15

Kernel source

aaaaf5c09e19.py29 lines
#!POPCORN leaderboard prefixsum_v2
import torch
from task import input_t, output_t

CUDA_SOURCE = r"""
// naive serial prefix sum — one thread does all the work
__global__ void solution(const float* __restrict__ input, float* __restrict__ output, int n) {
    if (threadIdx.x == 0 && blockIdx.x == 0) {
        double acc = 0.0;
        for (int i = 0; i < n; i++) {
            acc += (double)input[i];
            output[i] = (float)acc;
        }
    }
}

"""

_kernel = torch.cuda._compile_kernel(CUDA_SOURCE, "solution")

def custom_kernel(data: input_t) -> output_t:
    x, y = data
    n = x.numel()
    # adjust grid/block to match your kernel
    threads = (256, 1, 1)
    blocks = ((n + 255) // 256, 1, 1)
    _kernel(blocks, threads, (x, y, n))
    return y
scrolls · 29 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON