submission 66197
vinu1729 · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 76 lines, June 9 Researcher Reciprocity License v1.0.
test.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectoradd-v2-66197?include=source"interfacepython
Compatibility
measured onNVIDIA H100
declared hardwareNVIDIA H100
architecturessm_90
dtypesfp16
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:3e7f12c403ebeda6f0d9d63d496d9e6a3dd46482c1b5d81ec4cecc4505405a16
license declaredunknown
license concludedunknown
authorsvinu1729
imported2026-08-15
Kernel source
test.py76 lines
import torch
import triton
import triton.language as tl
@triton.jit
def vectoradd_kernel(
a_ptr, # Pointer to first input tensor
b_ptr, # Pointer to second input tensor
c_ptr, # Pointer to output tensor
N, # Total number of elements
BLOCK_SIZE: tl.constexpr, # Number of elements per block
):
"""
Optimized float16 vector addition kernel.
Each program processes BLOCK_SIZE elements.
"""
# Get program ID
pid = tl.program_id(0)
# Compute block start offset
block_start = pid * BLOCK_SIZE
# Generate offsets for this block
offsets = block_start + tl.arange(0, BLOCK_SIZE)
# Create mask for boundary checking (handles non-multiple of BLOCK_SIZE)
mask = offsets < N
# Load data with vectorized operations (float16)
# Use eviction policy for better cache utilization
a = tl.load(a_ptr + offsets, mask=mask, other=0.0, eviction_policy='evict_first')
b = tl.load(b_ptr + offsets, mask=mask, other=0.0, eviction_policy='evict_first')
# Perform addition
c = a + b
# Store result with vectorized write
tl.store(c_ptr + offsets, c, mask=mask)
def custom_kernel(data):
"""
Main kernel function that launches the Triton kernel.
Args:
data: Tuple of (A, B, C) where A and B are input tensors,
and C is the pre-allocated output tensor.
Returns:
C: Output tensor containing element-wise sum
"""
A, B, C = data
# Ensure inputs are contiguous for optimal memory access
A = A.contiguous()
B = B.contiguous()
# Get tensor dimensions
N = A.numel()
# Choose optimal block size based on GPU architecture
# Larger block sizes for better memory coalescing and fewer kernel launches
# Powers of 2 work best for GPU memory systems
BLOCK_SIZE = 1024 # Optimal for most modern GPUs (A100, H100, B200)
# Calculate grid size (number of programs to launch)
grid = lambda meta: (triton.cdiv(N, meta['BLOCK_SIZE']),)
# Launch kernel
vectoradd_kernel[grid](
A, B, C, N,
BLOCK_SIZE=BLOCK_SIZE,
)
return C
scrolls · 76 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Changes from previous submission
Against this author's previous submission submission 66193.
Best evidence level for this revision: reported
JSON