Skip to content
KernelIndex
Search⌘K

submission 68960

cdtmc · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 17 lines, June 9 Researcher Reciprocity License v1.0.

vectoradd1.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectoradd-v2-68960?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp16

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
FP16 vector additionsuite of 5 cases
NVIDIA B200
235.6µs
#20 of 66
2025-11-10

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:8bc497ed6dd8f59c686b9f55095755bb6490ac9791ff923e0885b714c53e5a4f
license declaredunknown
license concludedunknown
authorscdtmc
imported2026-08-15

Kernel source

vectoradd1.py17 lines
#!POPCORN leaderboard vectoradd_v2
#!POPCORN gpus B200 H100 A100 L4
from task import input_t, output_t
import torch


def add(x, y):
    return x + y


add_ = torch.compile(add)


def custom_kernel(data: input_t) -> output_t:
    d_in1, d_in2, d_out = data
    return add_(d_in1, d_in2)

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Changes from previous submission

Against this author's previous submission submission 68915.

- import triton
- import triton.language as tl
+ #!POPCORN leaderboard vectoradd_v2
+ #!POPCORN gpus B200 H100 A100 L4
+ from task import input_t, output_t
+ import torch
- @triton.autotune(
- configs=[
- # ------- Tier 1: good for L4 / small tensors -------
- triton.Config(kwargs={"BLOCK_SIZE": 256}, num_warps=2, num_stages=1),
- triton.Config(kwargs={"BLOCK_SIZE": 512}, num_warps=4, num_stages=1),
- # ------- Tier 2: A100 general sweet spot -------
- triton.Config(kwargs={"BLOCK_SIZE": 1024}, num_warps=4, num_stages=2),
- triton.Config(kwargs={"BLOCK_SIZE": 2048}, num_warps=8, num_stages=2),
- # ------- Tier 3: H100 / Blackwell optimized -------
- # Large blocks + pipelined loads
- triton.Config(kwargs={"BLOCK_SIZE": 4096}, num_warps=8, num_stages=3),
- triton.Config(kwargs={"BLOCK_SIZE": 8192}, num_warps=8, num_stages=4),
- ],
- key=["n_elements"], # autotune per problem size
- )
- @triton.jit
- def _kernel(x_ptr, y_ptr, output_ptr, n_elements, BLOCK_SIZE: tl.constexpr):
- pid = tl.program_id(axis=0)
- block_start = pid * BLOCK_SIZE
- offsets = block_start + tl.arange(0, BLOCK_SIZE)
- mask = offsets < n_elements
+ def add(x, y):
+ return x + y
- x = tl.load(x_ptr + offsets, mask=mask)
- y = tl.load(y_ptr + offsets, mask=mask)
- tl.store(output_ptr + offsets, x + y, mask=mask)
+ add_ = torch.compile(add)
- def custom_kernel(data):
- x, y, output = data
- n_elements = output.numel()
- grid = lambda meta: (triton.cdiv(n_elements, meta["BLOCK_SIZE"]),)
- _kernel[grid](x, y, output, n_elements)
- return output
+ def custom_kernel(data: input_t) -> output_t:
+ d_in1, d_in2, d_out = data
+ return add_(d_in1, d_in2)
scrolls · 48 diff lines total

Best evidence level for this revision: reported

JSON