submission 489093
jackkhuu · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 96 lines, June 9 Researcher Reciprocity License v1.0.
vectoradd_py_H100_claude-opus-4.5_ka_submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectoradd-v2-489093?include=source"interfacepython
Compatibility
measured onNVIDIA H100
declared hardwareNVIDIA H100
architecturessm_90
dtypesfp16
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:9ab7a6511f64435aeb0f4b398a379323466f2cdf11f693cf26f828a209736038
license declaredunknown
license concludedunknown
authorsjackkhuu
imported2026-08-15
Kernel source
vectoradd_py_H100_claude-opus-4.5_ka_submission.py96 lines
import triton
import triton.language as tl
import torch
@triton.jit
def _add_kernel(
a_ptr, # Pointer to input tensor A
b_ptr, # Pointer to input tensor B
c_ptr, # Pointer to output tensor C
n_elements, # Total number of elements
BLOCK_SIZE: tl.constexpr, # Number of elements per block
):
"""
Triton kernel for element-wise addition of two tensors.
Computes C = A + B for each element.
"""
# Get the program ID (which block we're in)
pid = tl.program_id(axis=0)
# Calculate the starting offset for this block
block_start = pid * BLOCK_SIZE
# Create offsets for elements this block will process
offsets = block_start + tl.arange(0, BLOCK_SIZE)
# Create mask to handle boundary conditions (last block may be partial)
mask = offsets < n_elements
# Load elements from A and B with masking
a = tl.load(a_ptr + offsets, mask=mask, other=0.0)
b = tl.load(b_ptr + offsets, mask=mask, other=0.0)
# Perform element-wise addition
c = a + b
# Store the result to C with masking
tl.store(c_ptr + offsets, c, mask=mask)
def kernel_function(A: torch.Tensor, B: torch.Tensor, C: torch.Tensor) -> torch.Tensor:
"""
Wrapper function for element-wise addition of two float16 tensors.
This is a single-operation kernel (no fusion needed) that computes C = A + B.
Args:
A: Input tensor of shape (size, size), dtype float16
B: Input tensor of shape (size, size), dtype float16
C: Output tensor of shape (size, size), dtype float16 (pre-allocated)
Returns:
C: The output tensor containing A + B
"""
# Validate inputs
assert A.is_cuda and B.is_cuda and C.is_cuda, "All tensors must be on CUDA"
assert A.shape == B.shape == C.shape, "All tensors must have the same shape"
assert A.is_contiguous() and B.is_contiguous() and C.is_contiguous(), "Tensors must be contiguous"
# Total number of elements to process
n_elements = A.numel()
# Choose block size (power of 2 for efficiency)
BLOCK_SIZE = 1024
# Calculate grid size (number of blocks needed)
grid = (triton.cdiv(n_elements, BLOCK_SIZE),)
# Launch the Triton kernel
_add_kernel[grid](
A, # Pointer to A
B, # Pointer to B
C, # Pointer to C
n_elements, # Total elements
BLOCK_SIZE=BLOCK_SIZE, # Compile-time constant
)
return C
import inspect
def custom_kernel(input):
sig = inspect.signature(kernel_function)
num_params = len(sig.parameters)
if len(input) == num_params:
return kernel_function(*input)
return kernel_function(input)
# Ensure deterministic cuBLAS.
import os
if os.environ.get("CUBLAS_WORKSPACE_CONFIG", "") not in (":4096:8", ":16:8"):
os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"
scrolls · 96 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON