Skip to content
KernelIndex
Search⌘K

submission 772562

ddwinterdd · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 59 lines, June 9 Researcher Reciprocity License v1.0.

submission_1_3.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectoradd-v2-772562?include=source"
interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp16

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
FP16 vector additionsuite of 5 cases
NVIDIA A100
896.0µs
#10 of 87
2026-04-16

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:8a67247a3d41009056dcdb183cbc60a67f6853732e620e37293a96f35867e6a7
license declaredunknown
license concludedunknown
authorsddwinterdd
imported2026-08-15

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

autotune@triton.autotune(
num-warps = 16triton.Config({}, num_warps=16),

Kernel source

submission_1_3.py59 lines
import triton
import triton.language as tl
import torch
from typing import Tuple

# Reference input/output type aliases (same as in reference.py)
# input_t = Tuple[torch.Tensor, torch.Tensor, torch.Tensor]
# output_t = torch.Tensor

@triton.autotune(
    configs=[
        triton.Config({}, num_warps=16),
    ],
    key=['N'],
)
@triton.jit
def add_kernel(A, B, C, N, BLOCK_SIZE: tl.constexpr):
    """Element‑wise addition of two 1‑D flattened tensors.

    A, B, C are pointers to the start of the tensors in device memory.
    N is the total number of elements.
    BLOCK_SIZE is the number of elements processed per program.
    """
    pid = tl.program_id(0)
    offs = pid * BLOCK_SIZE + tl.arange(0, BLOCK_SIZE)
    mask = offs < N

    a = tl.load(A + offs, mask=mask)
    b = tl.load(B + offs, mask=mask)
    c = a + b
    tl.store(C + offs, c, mask=mask)


def custom_kernel(data: Tuple[torch.Tensor, torch.Tensor, torch.Tensor]) -> torch.Tensor:
    """Entry point matching the reference signature.

    Args:
        data: (A, B, output) where A, B, output are torch.float16 CUDA tensors of shape (N, N).
    Returns:
        The output tensor containing A + B.
    """
    A, B, out = data
    assert A.is_cuda and B.is_cuda and out.is_cuda, "All tensors must be on CUDA"
    assert A.dtype == torch.float16 and B.dtype == torch.float16 and out.dtype == torch.float16, "All tensors must be float16"
    assert A.shape == B.shape == out.shape, "All tensors must have the same shape"

    # Flatten tensors for a 1‑D launch
    A_flat = A.view(-1)
    B_flat = B.view(-1)
    out_flat = out.view(-1)
    total_elements = A_flat.numel()

    # Choose a block size that is a power of two and fits into shared memory
    BLOCK_SIZE = 4096
    grid = lambda meta: ( (total_elements + meta['BLOCK_SIZE'] - 1) // meta['BLOCK_SIZE'], )

    add_kernel[grid](A_flat, B_flat, out_flat, total_elements, BLOCK_SIZE=BLOCK_SIZE)
    return out
scrolls · 59 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON