submission 67559
Nick · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 84 lines, June 9 Researcher Reciprocity License v1.0.
vectoradd2.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectoradd-v2-67559?include=source"interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp16
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:3add3c7a352302f060afa4d69b148662dca93a4e7daf5968dc3dddb33b2518af
license declaredunknown
license concludedunknown
authorsNick
imported2026-08-15
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
num-warps = 8
num_warps = 8stages = 1
num_stages=1,Kernel source
vectoradd2.py84 lines
from utils import make_match_reference, DeterministicContext
import torch
from task import input_t, output_t
import triton
import triton.language as tl
# Fixed optimal config - no autotuning variance
@triton.jit
def vecadd_fp16_kernel(A, B, C, N, BLOCK_SIZE: tl.constexpr):
pid = tl.program_id(0)
offs = pid * BLOCK_SIZE + tl.arange(0, BLOCK_SIZE)
mask = offs < N
# Load
a = tl.load(A + offs, mask=mask, other=0.0)
b = tl.load(B + offs, mask=mask, other=0.0)
# Compute
c = a + b
# Store
tl.store(C + offs, c, mask=mask)
def triton_vecadd(A, B, C):
N = A.numel()
# Fixed optimal config based on your 235us result
# Tune BLOCK_SIZE based on what worked best in autotuning
BLOCK_SIZE = 2048 # Start with this, adjust based on your best run
num_warps = 8
grid = (triton.cdiv(N, BLOCK_SIZE),)
vecadd_fp16_kernel[grid](
A, B, C, N,
BLOCK_SIZE=BLOCK_SIZE,
num_warps=num_warps,
num_stages=1,
)
return C
def ref_kernel(data: input_t) -> output_t:
"""
Reference implementation of vector addition using PyTorch.
Args:
data: Tuple of tensors [A, B, output] to be added.
Returns:
Tensor containing element-wise sums.
"""
with DeterministicContext():
A, B, output = data
output[...] = A + B
return output
def generate_input(size: int, seed: int) -> input_t:
"""
Generates random input tensors of specified shapes.
Returns:
Tuple of tensors [A, B, C] to be added.
"""
gen = torch.Generator(device="cuda")
gen.manual_seed(seed)
A = torch.randn(
size, size, device="cuda", dtype=torch.float16, generator=gen
).contiguous()
B = torch.randn(
size, size, device="cuda", dtype=torch.float16, generator=gen
).contiguous()
C = torch.empty(size, size, device="cuda", dtype=torch.float16).contiguous()
return A, B, C
def custom_kernel(data: input_t) -> output_t:
"""Fixed optimal Triton config - no autotuning variance"""
with DeterministicContext():
A, B, C = data
return triton_vecadd(A, B, C)
check_implementation = make_match_reference(ref_kernel)
scrolls · 84 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Changes from previous submission
Against this author's previous submission submission 67555.
⋯ 27 unchanged lines# Fixed optimal config based on your 235us result# Tune BLOCK_SIZE based on what worked best in autotuning- BLOCK_SIZE = 4096 # Start with this, adjust based on your best run- num_warps = 16+ BLOCK_SIZE = 2048 # Start with this, adjust based on your best run+ num_warps = 8grid = (triton.cdiv(N, BLOCK_SIZE),)vecadd_fp16_kernel[grid](
Best evidence level for this revision: reported
JSON