submission 68960
cdtmc · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 17 lines, June 9 Researcher Reciprocity License v1.0.
vectoradd1.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectoradd-v2-68960?include=source"interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp16
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:8bc497ed6dd8f59c686b9f55095755bb6490ac9791ff923e0885b714c53e5a4f
license declaredunknown
license concludedunknown
authorscdtmc
imported2026-08-15
Kernel source
vectoradd1.py17 lines
#!POPCORN leaderboard vectoradd_v2
#!POPCORN gpus B200 H100 A100 L4
from task import input_t, output_t
import torch
def add(x, y):
return x + y
add_ = torch.compile(add)
def custom_kernel(data: input_t) -> output_t:
d_in1, d_in2, d_out = data
return add_(d_in1, d_in2)
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Changes from previous submission
Against this author's previous submission submission 68915.
- import triton- import triton.language as tl+ #!POPCORN leaderboard vectoradd_v2+ #!POPCORN gpus B200 H100 A100 L4+ from task import input_t, output_t+ import torch- @triton.autotune(- configs=[- # ------- Tier 1: good for L4 / small tensors -------- triton.Config(kwargs={"BLOCK_SIZE": 256}, num_warps=2, num_stages=1),- triton.Config(kwargs={"BLOCK_SIZE": 512}, num_warps=4, num_stages=1),- # ------- Tier 2: A100 general sweet spot -------- triton.Config(kwargs={"BLOCK_SIZE": 1024}, num_warps=4, num_stages=2),- triton.Config(kwargs={"BLOCK_SIZE": 2048}, num_warps=8, num_stages=2),- # ------- Tier 3: H100 / Blackwell optimized -------- # Large blocks + pipelined loads- triton.Config(kwargs={"BLOCK_SIZE": 4096}, num_warps=8, num_stages=3),- triton.Config(kwargs={"BLOCK_SIZE": 8192}, num_warps=8, num_stages=4),- ],- key=["n_elements"], # autotune per problem size- )- @triton.jit- def _kernel(x_ptr, y_ptr, output_ptr, n_elements, BLOCK_SIZE: tl.constexpr):- pid = tl.program_id(axis=0)- block_start = pid * BLOCK_SIZE- offsets = block_start + tl.arange(0, BLOCK_SIZE)- mask = offsets < n_elements+ def add(x, y):+ return x + y- x = tl.load(x_ptr + offsets, mask=mask)- y = tl.load(y_ptr + offsets, mask=mask)- tl.store(output_ptr + offsets, x + y, mask=mask)+ add_ = torch.compile(add)- def custom_kernel(data):- x, y, output = data- n_elements = output.numel()- grid = lambda meta: (triton.cdiv(n_elements, meta["BLOCK_SIZE"]),)- _kernel[grid](x, y, output, n_elements)- return output+ def custom_kernel(data: input_t) -> output_t:+ d_in1, d_in2, d_out = data+ return add_(d_in1, d_in2)
scrolls · 48 diff lines total
Best evidence level for this revision: reported
JSON