submission 638685
sikuan · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 51 lines, June 9 Researcher Reciprocity License v1.0.
submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectorsum-v2-638685?include=source"interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp32
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:8af739382d919ce950e4e31a6454fd852077fdc8b6600ed967fa34747386d915
license declaredunknown
license concludedunknown
authorssikuan
imported2026-08-15
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
autotune
@triton.autotune(num-warps = 4
triton.Config({'BS': 1024}, num_warps=4),persistent-kernel
nprog = tl.num_programs(0)Kernel source
submission.py51 lines
import torch
import triton
import triton.language as tl
from task import input_t, output_t
_TRITON_THRESHOLD = 20_000_000
@triton.autotune(
configs=[
triton.Config({'BS': 1024}, num_warps=4),
triton.Config({'BS': 2048}, num_warps=4),
triton.Config({'BS': 2048}, num_warps=8),
triton.Config({'BS': 4096}, num_warps=8),
],
key=['N'],
)
@triton.jit
def _sum_pass1(x_ptr, p_ptr, N, BS: tl.constexpr):
pid = tl.program_id(0)
nprog = tl.num_programs(0)
acc = tl.zeros([BS], dtype=tl.float32)
for off in range(pid * BS, N, nprog * BS):
idx = off + tl.arange(0, BS)
acc += tl.load(x_ptr + idx, mask=idx < N, other=0.0)
tl.store(p_ptr + pid, tl.sum(acc, axis=0))
@triton.jit
def _final(p_ptr, o_ptr, M: tl.constexpr):
v = tl.load(p_ptr + tl.arange(0, M)).to(tl.float64)
tl.store(o_ptr, tl.sum(v, axis=0).to(tl.float32))
_NB = 256
_partial = None
def custom_kernel(data: input_t) -> output_t:
global _partial
input_tensor, output_tensor = data
n = input_tensor.numel()
if n <= _TRITON_THRESHOLD:
torch.sum(input_tensor, dim=0, keepdim=True, out=output_tensor)
else:
if _partial is None:
_partial = torch.empty(_NB, device=input_tensor.device, dtype=torch.float32)
_sum_pass1[(_NB,)](input_tensor, _partial, n)
_final[(1,)](_partial, output_tensor, M=_NB, num_warps=4)
return output_tensor[0]
scrolls · 51 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON