submission 545321
rajesh0042 · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 23 lines, June 9 Researcher Reciprocity License v1.0.
vectorsum_v4.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectorsum-v2-545321?include=source"interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp32
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:a58fb63be555dcead86692f7cc69313092e95109af4f4acc46b3ea32abbba94e
license declaredunknown
license concludedunknown
authorsrajesh0042
imported2026-08-15
Kernel source
vectorsum_v4.py23 lines
import os
os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"
import torch
from task import input_t, output_t
# Vectorsum: try to minimize overhead
# The reference does data.to(float64).sum().to(float32)
# Let's see if we can use torch.dot trick or kahan summation
# Pre-warm CUDA context
def _warmup():
for size in [1024, 4096, 16384]:
x = torch.randn(size, device='cuda', dtype=torch.float32)
x.double().sum().float()
torch.cuda.synchronize()
_warmup()
def custom_kernel(data: input_t) -> output_t:
data, output = data
return data.double().sum().float()
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Changes from previous submission
Against this author's previous submission submission 545126.
⋯ 1 unchanged linesos.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"import torch- import triton- import triton.language as tlfrom task import input_t, output_t- # Two-level reduction: GPU blocks reduce locally, then CPU sums partials- # Use float64 accumulation for accuracy matching reference+ # Vectorsum: try to minimize overhead+ # The reference does data.to(float64).sum().to(float32)+ # Let's see if we can use torch.dot trick or kahan summation- @triton.jit- def sum_reduce_kernel(- x_ptr, partial_ptr, n_elements,- BLOCK_SIZE: tl.constexpr,- ):- pid = tl.program_id(0)- offsets = pid * BLOCK_SIZE + tl.arange(0, BLOCK_SIZE)- mask = offsets < n_elements- # Load as float32, accumulate in float64- x = tl.load(x_ptr + offsets, mask=mask, other=0.0).to(tl.float64)- block_sum = tl.sum(x, axis=0)- tl.store(partial_ptr + pid, block_sum)+ # Pre-warm CUDA context+ def _warmup():+ for size in [1024, 4096, 16384]:+ x = torch.randn(size, device='cuda', dtype=torch.float32)+ x.double().sum().float()+ torch.cuda.synchronize()- @triton.jit- def sum_final_kernel(- partial_ptr, out_ptr, n_partials,- BLOCK_SIZE: tl.constexpr,- ):- offsets = tl.arange(0, BLOCK_SIZE)- mask = offsets < n_partials- x = tl.load(partial_ptr + offsets, mask=mask, other=0.0)- total = tl.sum(x, axis=0)- tl.store(out_ptr, total.to(tl.float32))+ _warmup()def custom_kernel(data: input_t) -> output_t:data, output = data- n = data.numel()- BLOCK_SIZE = 4096- n_blocks = (n + BLOCK_SIZE - 1) // BLOCK_SIZE- partial = torch.empty(n_blocks, device=data.device, dtype=torch.float64)- sum_reduce_kernel[(n_blocks,)](data, partial, n, BLOCK_SIZE=BLOCK_SIZE)-- # Second level reduction on GPU if too many partials- if n_blocks <= 65536:- # Use power-of-2 BLOCK_SIZE for final reduction- FINAL_BS = 1- while FINAL_BS < n_blocks:- FINAL_BS *= 2- if FINAL_BS > 65536:- FINAL_BS = 65536- out_scalar = torch.empty(1, device=data.device, dtype=torch.float32)- sum_final_kernel[(1,)](partial, out_scalar, n_blocks, BLOCK_SIZE=FINAL_BS)- return out_scalar[0]- else:- return partial.sum().to(torch.float32)+ return data.double().sum().float()
scrolls · 67 diff lines total
Best evidence level for this revision: reported
JSON