Skip to content
KernelIndex
Search⌘K

submission 512883

JordanNanos · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 52 lines, June 9 Researcher Reciprocity License v1.0.

vectorsum_py_submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectorsum-v2-512883?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
Vector sum reductionsuite of 6 cases
NVIDIA B200
60.3µs
#61 of 88
2026-03-05

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:0f25d3893ef2a31939d339fea63f5d034d108c1e5ce1eb249ce14a82db70fc93
license declaredunknown
license concludedunknown
authorsJordanNanos
imported2026-08-15

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

num-warps = 16num_warps=16,

Kernel source

vectorsum_py_submission.py52 lines
import torch
import triton
import triton.language as tl
from task import input_t, output_t


@triton.jit
def reduce_kernel(
    x_ptr,
    output_ptr,
    N,
    BLOCK_SIZE: tl.constexpr,
    NUM_BLOCKS: tl.constexpr,
):
    pid = tl.program_id(0)
    acc = tl.zeros([BLOCK_SIZE], dtype=tl.float64)
    num_chunks = tl.cdiv(N, BLOCK_SIZE)

    chunk = pid
    while chunk < num_chunks:
        offsets = chunk * BLOCK_SIZE + tl.arange(0, BLOCK_SIZE)
        mask = offsets < N
        x = tl.load(x_ptr + offsets, mask=mask, other=0.0).to(tl.float64)
        acc = acc + x
        chunk += NUM_BLOCKS

    local_sum = tl.sum(acc, axis=0).to(tl.float32)
    tl.atomic_add(output_ptr, local_sum)


def custom_kernel(data: input_t) -> output_t:
    x, output = data
    N = x.numel()

    if N <= 4096:
        output[0] = x.to(torch.float64).sum().to(torch.float32)
        return output[0]

    output.zero_()

    NUM_BLOCKS = 512
    BLOCK_SIZE = 4096

    reduce_kernel[(NUM_BLOCKS,)](
        x, output, N,
        BLOCK_SIZE=BLOCK_SIZE,
        NUM_BLOCKS=NUM_BLOCKS,
        num_warps=16,
    )

    return output[0]
scrolls · 52 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON