Skip to content
KernelIndex
Search⌘K

submission 545321

rajesh0042 · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 23 lines, June 9 Researcher Reciprocity License v1.0.

vectorsum_v4.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectorsum-v2-545321?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
Vector sum reductionsuite of 6 cases
NVIDIA B200
292.4µs
#83 of 88
2026-03-13

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:a58fb63be555dcead86692f7cc69313092e95109af4f4acc46b3ea32abbba94e
license declaredunknown
license concludedunknown
authorsrajesh0042
imported2026-08-15

Kernel source

vectorsum_v4.py23 lines
import os
os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"

import torch
from task import input_t, output_t

# Vectorsum: try to minimize overhead
# The reference does data.to(float64).sum().to(float32)
# Let's see if we can use torch.dot trick or kahan summation

# Pre-warm CUDA context
def _warmup():
    for size in [1024, 4096, 16384]:
        x = torch.randn(size, device='cuda', dtype=torch.float32)
        x.double().sum().float()
    torch.cuda.synchronize()

_warmup()

def custom_kernel(data: input_t) -> output_t:
    data, output = data
    return data.double().sum().float()

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Changes from previous submission

Against this author's previous submission submission 545126.

⋯ 1 unchanged lines
os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"
import torch
- import triton
- import triton.language as tl
from task import input_t, output_t
- # Two-level reduction: GPU blocks reduce locally, then CPU sums partials
- # Use float64 accumulation for accuracy matching reference
+ # Vectorsum: try to minimize overhead
+ # The reference does data.to(float64).sum().to(float32)
+ # Let's see if we can use torch.dot trick or kahan summation
- @triton.jit
- def sum_reduce_kernel(
- x_ptr, partial_ptr, n_elements,
- BLOCK_SIZE: tl.constexpr,
- ):
- pid = tl.program_id(0)
- offsets = pid * BLOCK_SIZE + tl.arange(0, BLOCK_SIZE)
- mask = offsets < n_elements
- # Load as float32, accumulate in float64
- x = tl.load(x_ptr + offsets, mask=mask, other=0.0).to(tl.float64)
- block_sum = tl.sum(x, axis=0)
- tl.store(partial_ptr + pid, block_sum)
+ # Pre-warm CUDA context
+ def _warmup():
+ for size in [1024, 4096, 16384]:
+ x = torch.randn(size, device='cuda', dtype=torch.float32)
+ x.double().sum().float()
+ torch.cuda.synchronize()
- @triton.jit
- def sum_final_kernel(
- partial_ptr, out_ptr, n_partials,
- BLOCK_SIZE: tl.constexpr,
- ):
- offsets = tl.arange(0, BLOCK_SIZE)
- mask = offsets < n_partials
- x = tl.load(partial_ptr + offsets, mask=mask, other=0.0)
- total = tl.sum(x, axis=0)
- tl.store(out_ptr, total.to(tl.float32))
+ _warmup()
def custom_kernel(data: input_t) -> output_t:
data, output = data
- n = data.numel()
- BLOCK_SIZE = 4096
- n_blocks = (n + BLOCK_SIZE - 1) // BLOCK_SIZE
- partial = torch.empty(n_blocks, device=data.device, dtype=torch.float64)
- sum_reduce_kernel[(n_blocks,)](data, partial, n, BLOCK_SIZE=BLOCK_SIZE)
-
- # Second level reduction on GPU if too many partials
- if n_blocks <= 65536:
- # Use power-of-2 BLOCK_SIZE for final reduction
- FINAL_BS = 1
- while FINAL_BS < n_blocks:
- FINAL_BS *= 2
- if FINAL_BS > 65536:
- FINAL_BS = 65536
- out_scalar = torch.empty(1, device=data.device, dtype=torch.float32)
- sum_final_kernel[(1,)](partial, out_scalar, n_blocks, BLOCK_SIZE=FINAL_BS)
- return out_scalar[0]
- else:
- return partial.sum().to(torch.float32)
+ return data.double().sum().float()
scrolls · 67 diff lines total

Best evidence level for this revision: reported

JSON