Skip to content
KernelIndex
Search⌘K

submission 638685

sikuan · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 51 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectorsum-v2-638685?include=source"
interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
Vector sum reductionsuite of 6 cases
NVIDIA A100
150.7µs
#50 of 96
2026-03-26

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:8af739382d919ce950e4e31a6454fd852077fdc8b6600ed967fa34747386d915
license declaredunknown
license concludedunknown
authorssikuan
imported2026-08-15

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

autotune@triton.autotune(
num-warps = 4triton.Config({'BS': 1024}, num_warps=4),
persistent-kernelnprog = tl.num_programs(0)

Kernel source

submission.py51 lines
import torch
import triton
import triton.language as tl
from task import input_t, output_t

_TRITON_THRESHOLD = 20_000_000


@triton.autotune(
    configs=[
        triton.Config({'BS': 1024}, num_warps=4),
        triton.Config({'BS': 2048}, num_warps=4),
        triton.Config({'BS': 2048}, num_warps=8),
        triton.Config({'BS': 4096}, num_warps=8),
    ],
    key=['N'],
)
@triton.jit
def _sum_pass1(x_ptr, p_ptr, N, BS: tl.constexpr):
    pid = tl.program_id(0)
    nprog = tl.num_programs(0)
    acc = tl.zeros([BS], dtype=tl.float32)
    for off in range(pid * BS, N, nprog * BS):
        idx = off + tl.arange(0, BS)
        acc += tl.load(x_ptr + idx, mask=idx < N, other=0.0)
    tl.store(p_ptr + pid, tl.sum(acc, axis=0))


@triton.jit
def _final(p_ptr, o_ptr, M: tl.constexpr):
    v = tl.load(p_ptr + tl.arange(0, M)).to(tl.float64)
    tl.store(o_ptr, tl.sum(v, axis=0).to(tl.float32))


_NB = 256
_partial = None


def custom_kernel(data: input_t) -> output_t:
    global _partial
    input_tensor, output_tensor = data
    n = input_tensor.numel()
    if n <= _TRITON_THRESHOLD:
        torch.sum(input_tensor, dim=0, keepdim=True, out=output_tensor)
    else:
        if _partial is None:
            _partial = torch.empty(_NB, device=input_tensor.device, dtype=torch.float32)
        _sum_pass1[(_NB,)](input_tensor, _partial, n)
        _final[(1,)](_partial, output_tensor, M=_NB, num_warps=4)
    return output_tensor[0]
scrolls · 51 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON