submission 510371
iharryli · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 66 lines, June 9 Researcher Reciprocity License v1.0.
submission_atomic_chunked.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectorsum-v2-510371?include=source"interfacepython
Compatibility
measured onNVIDIA L4
declared hardwareNVIDIA L4
architecturessm_89
dtypesfp32
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:ee53c3390e82565753ae72377e22c4ee21054d631aea9de81fdbb89307e61132
license declaredunknown
license concludedunknown
authorsiharryli
imported2026-08-15
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
num-warps = 1
_zero_scalar[(1,)](out, num_warps=1, num_stages=1)stages = 1
_zero_scalar[(1,)](out, num_warps=1, num_stages=1)Kernel source
submission_atomic_chunked.py66 lines
import torch
import triton
import triton.language as tl
from task import input_t, output_t
@triton.jit
def _sum_atomic_chunked(
x_ptr,
out_ptr,
n_elements,
BLOCK: tl.constexpr,
ITERS: tl.constexpr,
):
pid = tl.program_id(0)
base = pid * BLOCK * ITERS
r = tl.arange(0, BLOCK)
acc = tl.zeros((), dtype=tl.float32)
for i in tl.static_range(0, ITERS):
offsets = base + i * BLOCK + r
mask = offsets < n_elements
x = tl.load(x_ptr + offsets, mask=mask, other=0.0)
acc += tl.sum(x, axis=0)
tl.atomic_add(out_ptr, acc)
@triton.jit
def _zero_scalar(out_ptr):
tl.store(out_ptr, 0.0)
def _pick_iters(n: int) -> int:
# Reduce atomic traffic as N grows (baseline does 1 atomic per 1024 elems).
if n >= 20_000_000:
return 16
if n >= 6_000_000:
return 8
return 4
def custom_kernel(data: input_t) -> output_t:
x, out = data
n = x.numel()
iters = _pick_iters(n)
BLOCK = 1024
grid = (triton.cdiv(n, BLOCK * iters),)
# Ensure correct accumulation for atomic reduction.
_zero_scalar[(1,)](out, num_warps=1, num_stages=1)
_sum_atomic_chunked[grid](
x,
out,
n,
BLOCK=BLOCK,
ITERS=iters,
num_warps=8,
num_stages=4,
)
return out[0]
scrolls · 66 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON