Skip to content
KernelIndex
Search⌘K

submission 780472

brianyu · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 41 lines, June 9 Researcher Reciprocity License v1.0.

vecadd2.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectoradd-v2-780472?include=source"
interfacepython
Compatibility
measured onNVIDIA L4
declared hardwareNVIDIA L4
architecturessm_89
dtypesfp16

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
FP16 vector additionsuite of 5 cases
NVIDIA L4
6.53ms
#5 of 26
2026-04-29

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:d2397a4a4bc8ddf44695db94ccb261eaaf01556038b318c38e75ad51d4bd0948
license declaredunknown
license concludedunknown
authorsbrianyu
imported2026-08-15

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

num-warps = 8_add_kernel[grid](A, B, output, BLOCK_SIZE=block_size, num_warps=8)

Kernel source

vecadd2.py41 lines
#!POPCORN leaderboard vectoradd_v2
#!POPCORN gpus L4

from task import input_t, output_t

import torch
import triton
import triton.language as tl


@triton.jit
def _add_kernel(
    a_ptr,
    b_ptr,
    out_ptr,
    BLOCK_SIZE: tl.constexpr,
):
    pid = tl.program_id(0)
    offsets = pid * BLOCK_SIZE + tl.arange(0, BLOCK_SIZE)
    offsets = tl.max_contiguous(tl.multiple_of(offsets, BLOCK_SIZE), BLOCK_SIZE)

    a = tl.load(a_ptr + offsets, eviction_policy="evict_first")
    b = tl.load(b_ptr + offsets, eviction_policy="evict_first")
    tl.store(out_ptr + offsets, a + b)


def custom_kernel(data: input_t) -> output_t:
    # vector-add-v2 passes an output buffer in the input tuple.  Reusing it
    # avoids a very visible allocation cost on the 8192/16384 cases.
    if len(data) == 3:
        A, B, output = data
    else:
        A, B = data
        output = torch.empty_like(A)

    n_elements = A.numel()
    block_size = 512
    grid = (triton.cdiv(n_elements, block_size),)
    _add_kernel[grid](A, B, output, BLOCK_SIZE=block_size, num_warps=8)
    return output
scrolls · 41 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON