submission 780039
Cookie 🍪 · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 69 lines, June 9 Researcher Reciprocity License v1.0.
vectoradd_v2_H100_claude-opus-4.6_ka_submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectoradd-v2-780039?include=source"interfacepython
Compatibility
measured onNVIDIA H100
declared hardwareNVIDIA H100
architecturessm_90
dtypesfp16
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:c770836962013ff9dffb5e59a251553574ece0a949189e5fc32edacd45070a89
license declaredunknown
license concludedunknown
authorsCookie 🍪
imported2026-08-15
Kernel source
vectoradd_v2_H100_claude-opus-4.6_ka_submission.py69 lines
import triton
import triton.language as tl
import torch
@triton.jit
def _vecadd_kernel(ptr_a, ptr_b, ptr_c, n_elements, BLOCK_SIZE: tl.constexpr):
"""Triton kernel for elementwise addition of two float16 tensors.
Fused stages: single elementwise add (a + b -> c).
No further fusion possible since this is a standalone vector addition.
"""
pid = tl.program_id(0)
block_start = pid * BLOCK_SIZE
offsets = block_start + tl.arange(0, BLOCK_SIZE)
mask = offsets < n_elements
a = tl.load(ptr_a + offsets, mask=mask)
b = tl.load(ptr_b + offsets, mask=mask)
c = a + b
tl.store(ptr_c + offsets, c, mask=mask)
def kernel_function(A: torch.Tensor, B: torch.Tensor, C: torch.Tensor) -> torch.Tensor:
"""Wrapper for float16 vector addition: C = A + B.
Args:
A: Input tensor of shape (N, N), dtype float16, on CUDA.
B: Input tensor of shape (N, N), dtype float16, on CUDA.
C: Output tensor of shape (N, N), dtype float16, on CUDA (written in-place).
Returns:
C tensor with result A + B.
"""
assert A.is_cuda and B.is_cuda and C.is_cuda
assert A.shape == B.shape == C.shape
# Flatten for 1D indexing - use contiguous views
a_flat = A.contiguous().view(-1)
b_flat = B.contiguous().view(-1)
c_flat = C.contiguous().view(-1)
n_elements = a_flat.numel()
BLOCK_SIZE = 1024
grid = (triton.cdiv(n_elements, BLOCK_SIZE),)
_vecadd_kernel[grid](a_flat, b_flat, c_flat, n_elements, BLOCK_SIZE=BLOCK_SIZE)
return C
import inspect
def custom_kernel(input):
sig = inspect.signature(kernel_function)
num_params = len(sig.parameters)
if len(input) == num_params:
return kernel_function(*input)
return kernel_function(input)
# Ensure deterministic cuBLAS.
import os
if os.environ.get("CUBLAS_WORKSPACE_CONFIG", "") not in (":4096:8", ":16:8"):
os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"
scrolls · 69 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON