Skip to content
KernelIndex
Search⌘K

submission 780039

Cookie 🍪 · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 69 lines, June 9 Researcher Reciprocity License v1.0.

vectoradd_v2_H100_claude-opus-4.6_ka_submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectoradd-v2-780039?include=source"
interfacepython
Compatibility
measured onNVIDIA H100
declared hardwareNVIDIA H100
architecturessm_90
dtypesfp16

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
FP16 vector additionsuite of 5 cases
NVIDIA H100
537.8µs
#38 of 44
2026-04-24

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:c770836962013ff9dffb5e59a251553574ece0a949189e5fc32edacd45070a89
license declaredunknown
license concludedunknown
authorsCookie 🍪
imported2026-08-15

Kernel source

vectoradd_v2_H100_claude-opus-4.6_ka_submission.py69 lines
import triton
import triton.language as tl
import torch


@triton.jit
def _vecadd_kernel(ptr_a, ptr_b, ptr_c, n_elements, BLOCK_SIZE: tl.constexpr):
    """Triton kernel for elementwise addition of two float16 tensors.
    
    Fused stages: single elementwise add (a + b -> c).
    No further fusion possible since this is a standalone vector addition.
    """
    pid = tl.program_id(0)
    block_start = pid * BLOCK_SIZE
    offsets = block_start + tl.arange(0, BLOCK_SIZE)
    mask = offsets < n_elements

    a = tl.load(ptr_a + offsets, mask=mask)
    b = tl.load(ptr_b + offsets, mask=mask)

    c = a + b

    tl.store(ptr_c + offsets, c, mask=mask)


def kernel_function(A: torch.Tensor, B: torch.Tensor, C: torch.Tensor) -> torch.Tensor:
    """Wrapper for float16 vector addition: C = A + B.

    Args:
        A: Input tensor of shape (N, N), dtype float16, on CUDA.
        B: Input tensor of shape (N, N), dtype float16, on CUDA.
        C: Output tensor of shape (N, N), dtype float16, on CUDA (written in-place).

    Returns:
        C tensor with result A + B.
    """
    assert A.is_cuda and B.is_cuda and C.is_cuda
    assert A.shape == B.shape == C.shape

    # Flatten for 1D indexing - use contiguous views
    a_flat = A.contiguous().view(-1)
    b_flat = B.contiguous().view(-1)
    c_flat = C.contiguous().view(-1)

    n_elements = a_flat.numel()
    BLOCK_SIZE = 1024
    grid = (triton.cdiv(n_elements, BLOCK_SIZE),)

    _vecadd_kernel[grid](a_flat, b_flat, c_flat, n_elements, BLOCK_SIZE=BLOCK_SIZE)

    return C

import inspect

def custom_kernel(input):
    sig = inspect.signature(kernel_function)
    num_params = len(sig.parameters)

    if len(input) == num_params:
        return kernel_function(*input)
    return kernel_function(input)


# Ensure deterministic cuBLAS.
import os
if os.environ.get("CUBLAS_WORKSPACE_CONFIG", "") not in (":4096:8", ":16:8"):
    os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"

scrolls · 69 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON