submission 614137
dannywillowliu-uchi · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 46 lines, June 9 Researcher Reciprocity License v1.0.
submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-prefixsum-v2-614137?include=source"interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp32
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:543a6f04407e41efdaf6bf70dca18e7b779e59def58782e3276f21b5957422c7
license declaredunknown
license concludedunknown
authorsdannywillowliu-uchi
imported2026-08-15
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
num-warps = 4
_reduce[(nb,)](inp, _bsums, n, BS=BS, num_warps=4)Kernel source
submission.py46 lines
import os
os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"
import torch
import triton
import triton.language as tl
from task import input_t, output_t
BS = 1024
@triton.jit
def _reduce(inp_ptr, bsums_ptr, n, BS: tl.constexpr):
pid = tl.program_id(0)
offs = pid * BS + tl.arange(0, BS)
m = offs < n
x = tl.load(inp_ptr + offs, mask=m, other=0.0)
tl.store(bsums_ptr + pid, tl.sum(x))
@triton.jit
def _scan_add(inp_ptr, out_ptr, ps_ptr, n, BS: tl.constexpr):
pid = tl.program_id(0)
offs = pid * BS + tl.arange(0, BS)
m = offs < n
x = tl.load(inp_ptr + offs, mask=m, other=0.0)
s = tl.cumsum(x)
if pid > 0:
p = tl.load(ps_ptr + pid - 1)
else:
p = 0.0
tl.store(out_ptr + offs, s + p, mask=m)
_bsums = None
_psums = None
def custom_kernel(data: input_t) -> output_t:
global _bsums, _psums
inp, out = data
n = inp.numel()
nb = (n + BS - 1) // BS
if _bsums is None or _bsums.numel() < nb:
_bsums = torch.empty(nb, device="cuda", dtype=torch.float32)
_psums = torch.empty(nb, device="cuda", dtype=torch.float32)
_reduce[(nb,)](inp, _bsums, n, BS=BS, num_warps=4)
torch.cumsum(_bsums[:nb], dim=0, out=_psums[:nb])
_scan_add[(nb,)](inp, out, _psums, n, BS=BS, num_warps=4)
return out
scrolls · 46 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Changes from previous submission
Against this author's previous submission submission 612455.
Best evidence level for this revision: reported
JSON