Skip to content
KernelIndex
Search⌘K

submission 780407

Kazim · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 43 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectorsum-v2-780407?include=source"
interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
Vector sum reductionsuite of 6 cases
NVIDIA A100
147.0µs
#38 of 96
2026-04-27

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:2454ab7f7c19131a510ed1334fbb9c72021308459884b9839f0380754eee1a3b
license declaredunknown
license concludedunknown
authorsKazim
imported2026-08-15

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

num-warps = 16NUM_WARPS = 16
persistent-kerneldef persistent_sum_kernel(x_ptr, out_ptr, n_elements, BLOCK_SIZE: tl.constexpr):

Kernel source

submission.py43 lines
#!POPCORN leaderboard vectorsum_v2
#!POPCORN gpu A100

import torch
import triton
import triton.language as tl

# A100: 108 SMs, 2048 max threads/SM.
# 16 warps (512 threads/block) × 4 blocks/SM = 64 warps/SM = 100% occupancy.
# This saturates HBM bandwidth so the GPU never stalls waiting for data.
NUM_SMS = 108
GRID = NUM_SMS * 4   # 432 blocks
BLOCK_SIZE = 8192
NUM_WARPS = 16


@triton.jit
def persistent_sum_kernel(x_ptr, out_ptr, n_elements, BLOCK_SIZE: tl.constexpr):
    pid = tl.program_id(0)
    num_programs = tl.num_programs(0)

    acc = tl.zeros([BLOCK_SIZE], dtype=tl.float32)
    block_start = pid * BLOCK_SIZE
    stride = num_programs * BLOCK_SIZE

    while block_start < n_elements:
        offsets = block_start + tl.arange(0, BLOCK_SIZE)
        mask = offsets < n_elements
        x = tl.load(x_ptr + offsets, mask=mask, other=0.0,
                     eviction_policy='evict_first').to(tl.float32)
        acc += x
        block_start += stride

    tl.atomic_add(out_ptr, tl.sum(acc, axis=0))


def custom_kernel(data) -> torch.Tensor:
    x = data[0].cuda()
    out = torch.zeros(1, dtype=torch.float32, device=x.device)
    persistent_sum_kernel[(GRID,)](x, out, x.numel(),
                                   BLOCK_SIZE=BLOCK_SIZE, num_warps=NUM_WARPS)
    return out[0]
scrolls · 43 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON