submission 67714
Pritam · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 46 lines, June 9 Researcher Reciprocity License v1.0.
submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectoradd-v2-67714?include=source"interfacepython
Compatibility
measured onNVIDIA H100
declared hardwareNVIDIA H100
architecturessm_90
dtypesfp16
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:fa419b80e068c57ae2067a43d754bc1bb80d165ff636570fb418064d189cba66
license declaredunknown
license concludedunknown
authorsPritam
imported2026-08-15
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
autotune
@triton.autotune(num-warps = 8
triton.Config({"BLOCK_SIZE": 8192}, num_warps=8, num_stages=6),stages = 6
triton.Config({"BLOCK_SIZE": 8192}, num_warps=8, num_stages=6),Kernel source
submission.py46 lines
#!POPCORN leaderboard vectoradd_v2
import torch, triton, triton.language as tl
from task import input_t, output_t
@triton.autotune(
configs=[
triton.Config({"BLOCK_SIZE": 8192}, num_warps=8, num_stages=6),
triton.Config({"BLOCK_SIZE": 4096}, num_warps=8, num_stages=5),
triton.Config({"BLOCK_SIZE": 16384}, num_warps=8, num_stages=6),
],
key=[],
)
@triton.jit
def vecadd_kernel(A_ptr, B_ptr, C_ptr, N, BLOCK_SIZE: tl.constexpr):
pid = tl.program_id(0)
SPAN: tl.constexpr = BLOCK_SIZE * 4 # power-of-two span → OK for arange
offs = pid * SPAN + tl.arange(0, SPAN)
mask = offs < N
# Encourage coalesced 128B vector loads
tl.multiple_of(offs, 128)
tl.max_contiguous(offs, SPAN)
# Fast streaming load (L2 only), use fp16 math pipeline friendly pattern
a = tl.load(A_ptr + offs, mask=mask, other=0.0, cache_modifier=".cg")
b = tl.load(B_ptr + offs, mask=mask, other=0.0, cache_modifier=".cg")
c = a + b
# Avoid extra hazards — async commit path on Hopper is faster for large N
tl.store(C_ptr + offs, c, mask=mask)
def custom_kernel(data: input_t) -> output_t:
A, B, C = data
N = A.numel()
# ensure contiguous access (esp. for Popcorn harness copies)
if not A.is_contiguous(): A = A.contiguous()
if not B.is_contiguous(): B = B.contiguous()
grid = lambda META: (triton.cdiv(N, META["BLOCK_SIZE"] * 4),)
vecadd_kernel[grid](A, B, C, N)
return C
scrolls · 46 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON