submission 780407
Kazim · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 43 lines, June 9 Researcher Reciprocity License v1.0.
submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectorsum-v2-780407?include=source"interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp32
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:2454ab7f7c19131a510ed1334fbb9c72021308459884b9839f0380754eee1a3b
license declaredunknown
license concludedunknown
authorsKazim
imported2026-08-15
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
num-warps = 16
NUM_WARPS = 16persistent-kernel
def persistent_sum_kernel(x_ptr, out_ptr, n_elements, BLOCK_SIZE: tl.constexpr):Kernel source
submission.py43 lines
#!POPCORN leaderboard vectorsum_v2
#!POPCORN gpu A100
import torch
import triton
import triton.language as tl
# A100: 108 SMs, 2048 max threads/SM.
# 16 warps (512 threads/block) × 4 blocks/SM = 64 warps/SM = 100% occupancy.
# This saturates HBM bandwidth so the GPU never stalls waiting for data.
NUM_SMS = 108
GRID = NUM_SMS * 4 # 432 blocks
BLOCK_SIZE = 8192
NUM_WARPS = 16
@triton.jit
def persistent_sum_kernel(x_ptr, out_ptr, n_elements, BLOCK_SIZE: tl.constexpr):
pid = tl.program_id(0)
num_programs = tl.num_programs(0)
acc = tl.zeros([BLOCK_SIZE], dtype=tl.float32)
block_start = pid * BLOCK_SIZE
stride = num_programs * BLOCK_SIZE
while block_start < n_elements:
offsets = block_start + tl.arange(0, BLOCK_SIZE)
mask = offsets < n_elements
x = tl.load(x_ptr + offsets, mask=mask, other=0.0,
eviction_policy='evict_first').to(tl.float32)
acc += x
block_start += stride
tl.atomic_add(out_ptr, tl.sum(acc, axis=0))
def custom_kernel(data) -> torch.Tensor:
x = data[0].cuda()
out = torch.zeros(1, dtype=torch.float32, device=x.device)
persistent_sum_kernel[(GRID,)](x, out, x.numel(),
BLOCK_SIZE=BLOCK_SIZE, num_warps=NUM_WARPS)
return out[0]
scrolls · 43 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON