submission 776466
x3C49 · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 61 lines, June 9 Researcher Reciprocity License v1.0.
submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectorsum-v2-776466?include=source"interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp32
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:73a71384827ff6b698ccca53f4cc0dfcde3f3dcc1978a1ef646379e390ce4a0a
license declaredunknown
license concludedunknown
authorsx3C49
imported2026-08-15
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
num-warps = 16
num_warps=16,Kernel source
submission.py61 lines
import torch
import triton
import triton.language as tl
from task import input_t, output_t
GRID = 512
BLOCK = 8192 # Doubled block size → loop iterations halved (25 → 13)
@triton.jit
def _reduce_pass1(
data_ptr, partial_ptr, N,
BLOCK: tl.constexpr,
GRID: tl.constexpr,
):
pid = tl.program_id(0)
acc = tl.zeros((BLOCK,), dtype=tl.float32) # FP32 halves register use
start = pid * BLOCK
stride = GRID * BLOCK
while start + BLOCK <= N:
acc += tl.load(data_ptr + start + tl.arange(0, BLOCK)).to(tl.float32)
start += stride
if start < N:
offs = start + tl.arange(0, BLOCK)
acc += tl.load(data_ptr + offs, mask=offs < N, other=0.0).to(tl.float32)
tl.store(partial_ptr + pid, tl.sum(acc, axis=0))
@triton.jit
def _reduce_pass2(partial_ptr, out_ptr, GRID: tl.constexpr):
s = tl.sum(tl.load(partial_ptr + tl.arange(0, GRID)), axis=0)
tl.store(out_ptr, s.to(tl.float32))
_scratch: torch.Tensor | None = None
def custom_kernel(data: input_t) -> output_t:
global _scratch
input_tensor, output_tensor = data
N = input_tensor.numel()
if _scratch is None:
_scratch = torch.empty(GRID, device="cuda", dtype=torch.float32)
_reduce_pass1[(GRID,)](
input_tensor, _scratch, N,
BLOCK=BLOCK,
GRID=GRID,
num_warps=16,
)
_reduce_pass2[(1,)](
_scratch, output_tensor,
GRID=GRID,
num_warps=4,
)
return output_tensor.view([])scrolls · 61 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Changes from previous submission
Against this author's previous submission submission 776241.
⋯ 3 unchanged linesfrom task import input_t, output_tGRID = 512- BLOCK = 4096+ BLOCK = 8192 # Doubled block size → loop iterations halved (25 → 13)@triton.jit⋯ 3 unchanged linesGRID: tl.constexpr,):pid = tl.program_id(0)- acc = tl.zeros((BLOCK,), dtype=tl.float64)+ acc = tl.zeros((BLOCK,), dtype=tl.float32) # FP32 halves register usestart = pid * BLOCKstride = GRID * BLOCK- # ── hot path: full tiles, no mask ──────────────────────────────────────while start + BLOCK <= N:- acc += tl.load(data_ptr + start + tl.arange(0, BLOCK)).to(tl.float64)+ acc += tl.load(data_ptr + start + tl.arange(0, BLOCK)).to(tl.float32)start += stride- # ── tail: partial tile (not reached for any benchmark size) ────────────if start < N:offs = start + tl.arange(0, BLOCK)- acc += tl.load(data_ptr + offs, mask=offs < N, other=0.0).to(tl.float64)+ acc += tl.load(data_ptr + offs, mask=offs < N, other=0.0).to(tl.float32)tl.store(partial_ptr + pid, tl.sum(acc, axis=0))⋯ 13 unchanged linesN = input_tensor.numel()if _scratch is None:- _scratch = torch.empty(GRID, device="cuda", dtype=torch.float64)+ _scratch = torch.empty(GRID, device="cuda", dtype=torch.float32)_reduce_pass1[(GRID,)](input_tensor, _scratch, N,BLOCK=BLOCK,GRID=GRID,- num_warps=16, # 512 threads/block → 4 blocks/SM → 64 warps/SM = 100% occupancy+ num_warps=16,)_reduce_pass2[(1,)](_scratch, output_tensor,
scrolls · 48 diff lines total
Best evidence level for this revision: reported
JSON