submission 512774
mreso · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 49 lines, June 9 Researcher Reciprocity License v1.0.
submission_vectorsum_v2.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectorsum-v2-512774?include=source"interfacepython
Compatibility
measured onNVIDIA H100
declared hardwareNVIDIA H100
architecturessm_90
dtypesfp32
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:caa96fa064b86e6ad58b8030e532c5c4318270dc1a0e1787fda5ad689901b568
license declaredunknown
license concludedunknown
authorsmreso
imported2026-08-15
Kernel source
submission_vectorsum_v2.py49 lines
# submission_vectorsum_v2.py
# Sum reduction of a 1D float32 tensor → scalar.
# Interface: custom_kernel((input_tensor, output_tensor)) -> output_tensor
# input_tensor: (size,) float32
# output_tensor: (1,) float32 (pre-allocated, but reference returns 0-d)
#
# The reference returns data.to(float64).sum().to(float32) — a 0-d tensor.
# We must return a 0-d tensor too to pass the shape check in check_implementation.
# Float64 accumulators preserve accuracy for large arrays.
import torch
import triton
import triton.language as tl
from task import input_t, output_t
_BLOCK = 4096
_WARPS = 8
@triton.jit
def _partial_reduce(x_ptr, partial_ptr, N: int, BLOCK: tl.constexpr):
"""Reduce BLOCK float32 elements → one float64 partial sum."""
pid = tl.program_id(0)
offs = pid * BLOCK + tl.arange(0, BLOCK)
mask = offs < N
x = tl.load(x_ptr + offs, mask=mask, other=0.0).to(tl.float64)
tl.store(partial_ptr + pid, tl.sum(x, axis=0))
def custom_kernel(data: input_t) -> output_t:
x, out = data
N = x.numel()
if N == 0:
result = torch.zeros((), dtype=torch.float32, device=x.device)
out[0] = result
return result
P = triton.cdiv(N, _BLOCK)
partial = torch.empty(P, dtype=torch.float64, device=x.device)
_partial_reduce[(P,)](x, partial, N, BLOCK=_BLOCK, num_warps=_WARPS)
# Final reduction on P float64 partials → 0-d float32 scalar.
# The reference also returns a 0-d tensor, so we match that shape.
result = partial.sum().to(torch.float32) # 0-d tensor
out[0] = result
return result
scrolls · 49 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Changes from previous submission
Against this author's previous submission submission 512770.
Best evidence level for this revision: reported
JSON