Skip to content
KernelIndex
Search⌘K

submission 779742

x3C49 · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 46 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectorsum-v2-779742?include=source"
interfacepython
Compatibility
measured onNVIDIA H100
declared hardwareNVIDIA H100
architecturessm_90
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
Vector sum reductionsuite of 6 cases
NVIDIA H100
86.9µs
#22 of 37
2026-04-23

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:35263090feb7e511c482b2b2750459748975594c7ce062f64555a9d424b62649
license declaredunknown
license concludedunknown
authorsx3C49
imported2026-08-15

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

num-warps = 16num_warps=16,

Kernel source

submission.py46 lines
import torch
import triton
import triton.language as tl
from task import input_t, output_t
GRID = 512
BLOCK = 8192          # Doubled block size → loop iterations halved (25 → 13)
@triton.jit
def _reduce_pass1(
    data_ptr, partial_ptr, N,
    BLOCK: tl.constexpr,
    GRID: tl.constexpr,
):
    pid = tl.program_id(0)
    acc = tl.zeros((BLOCK,), dtype=tl.float32)   # FP32 halves register use
    start = pid * BLOCK
    stride = GRID * BLOCK
    while start + BLOCK <= N:
        acc += tl.load(data_ptr + start + tl.arange(0, BLOCK)).to(tl.float32)
        start += stride
    if start < N:
        offs = start + tl.arange(0, BLOCK)
        acc += tl.load(data_ptr + offs, mask=offs < N, other=0.0).to(tl.float32)
    tl.store(partial_ptr + pid, tl.sum(acc, axis=0))
@triton.jit
def _reduce_pass2(partial_ptr, out_ptr, GRID: tl.constexpr):
    s = tl.sum(tl.load(partial_ptr + tl.arange(0, GRID)), axis=0)
    tl.store(out_ptr, s.to(tl.float32))
_scratch: torch.Tensor | None = None
def custom_kernel(data: input_t) -> output_t:
    global _scratch
    input_tensor, output_tensor = data
    N = input_tensor.numel()
    if _scratch is None:
        _scratch = torch.empty(GRID, device="cuda", dtype=torch.float32)
    _reduce_pass1[(GRID,)](
        input_tensor, _scratch, N,
        BLOCK=BLOCK,
        GRID=GRID,
        num_warps=16,
    )
    _reduce_pass2[(1,)](
        _scratch, output_tensor,
        GRID=GRID,
        num_warps=4,
    )
    return output_tensor.view([])
scrolls · 46 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Changes from previous submission

Against this author's previous submission submission 776466.

⋯ 1 unchanged lines
import triton
import triton.language as tl
from task import input_t, output_t
-
GRID = 512
BLOCK = 8192 # Doubled block size → loop iterations halved (25 → 13)
-
-
@triton.jit
def _reduce_pass1(
data_ptr, partial_ptr, N,
⋯ 4 unchanged lines
acc = tl.zeros((BLOCK,), dtype=tl.float32) # FP32 halves register use
start = pid * BLOCK
stride = GRID * BLOCK
-
while start + BLOCK <= N:
acc += tl.load(data_ptr + start + tl.arange(0, BLOCK)).to(tl.float32)
start += stride
-
if start < N:
offs = start + tl.arange(0, BLOCK)
acc += tl.load(data_ptr + offs, mask=offs < N, other=0.0).to(tl.float32)
-
tl.store(partial_ptr + pid, tl.sum(acc, axis=0))
-
-
@triton.jit
def _reduce_pass2(partial_ptr, out_ptr, GRID: tl.constexpr):
s = tl.sum(tl.load(partial_ptr + tl.arange(0, GRID)), axis=0)
tl.store(out_ptr, s.to(tl.float32))
-
-
_scratch: torch.Tensor | None = None
-
-
def custom_kernel(data: input_t) -> output_t:
global _scratch
input_tensor, output_tensor = data
N = input_tensor.numel()
-
if _scratch is None:
_scratch = torch.empty(GRID, device="cuda", dtype=torch.float32)
-
_reduce_pass1[(GRID,)](
input_tensor, _scratch, N,
BLOCK=BLOCK,
⋯ 5 unchanged lines
GRID=GRID,
num_warps=4,
)
-
return output_tensor.view([])
No newline at end of file
scrolls · 55 diff lines total

Best evidence level for this revision: reported

JSON