Skip to content
KernelIndex
Search⌘K

submission 512774

mreso · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 49 lines, June 9 Researcher Reciprocity License v1.0.

submission_vectorsum_v2.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectorsum-v2-512774?include=source"
interfacepython
Compatibility
measured onNVIDIA H100
declared hardwareNVIDIA H100
architecturessm_90
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
Vector sum reductionsuite of 6 cases
NVIDIA H100
94.2µs
#34 of 37
2026-03-04

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:caa96fa064b86e6ad58b8030e532c5c4318270dc1a0e1787fda5ad689901b568
license declaredunknown
license concludedunknown
authorsmreso
imported2026-08-15

Kernel source

submission_vectorsum_v2.py49 lines
# submission_vectorsum_v2.py
# Sum reduction of a 1D float32 tensor → scalar.
# Interface: custom_kernel((input_tensor, output_tensor)) -> output_tensor
#   input_tensor:  (size,) float32
#   output_tensor: (1,)   float32 (pre-allocated, but reference returns 0-d)
#
# The reference returns data.to(float64).sum().to(float32) — a 0-d tensor.
# We must return a 0-d tensor too to pass the shape check in check_implementation.
# Float64 accumulators preserve accuracy for large arrays.

import torch
import triton
import triton.language as tl
from task import input_t, output_t

_BLOCK = 4096
_WARPS = 8


@triton.jit
def _partial_reduce(x_ptr, partial_ptr, N: int, BLOCK: tl.constexpr):
    """Reduce BLOCK float32 elements → one float64 partial sum."""
    pid  = tl.program_id(0)
    offs = pid * BLOCK + tl.arange(0, BLOCK)
    mask = offs < N
    x    = tl.load(x_ptr + offs, mask=mask, other=0.0).to(tl.float64)
    tl.store(partial_ptr + pid, tl.sum(x, axis=0))


def custom_kernel(data: input_t) -> output_t:
    x, out = data
    N = x.numel()

    if N == 0:
        result = torch.zeros((), dtype=torch.float32, device=x.device)
        out[0] = result
        return result

    P       = triton.cdiv(N, _BLOCK)
    partial = torch.empty(P, dtype=torch.float64, device=x.device)

    _partial_reduce[(P,)](x, partial, N, BLOCK=_BLOCK, num_warps=_WARPS)

    # Final reduction on P float64 partials → 0-d float32 scalar.
    # The reference also returns a 0-d tensor, so we match that shape.
    result = partial.sum().to(torch.float32)  # 0-d tensor
    out[0] = result
    return result
scrolls · 49 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Changes from previous submission

Against this author's previous submission submission 512770.

Best evidence level for this revision: reported

JSON