submission 648869
JNJYan · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 56 lines, June 9 Researcher Reciprocity License v1.0.
l4.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectoradd-v2-648869?include=source"interfacepython
Compatibility
measured onNVIDIA L4
declared hardwareNVIDIA L4
architecturessm_89
dtypesfp16
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:78467c8e8a4421da04eff2b247ad8e05f6da5317ee74499f9e0a5dd501382623
license declaredunknown
license concludedunknown
authorsJNJYan
imported2026-08-15
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
num-warps = 8
_vecadd_kernel[grid](A_1d, B_1d, O_1d, n_elements, BLOCK=1024, num_warps=8)Kernel source
l4.py56 lines
#!POPCORN leaderboard vectoradd_v2
#!POPCORN gpu L4
from task import input_t, output_t
try:
import triton as _triton
import triton.language as _tl
_TRITON_AVAILABLE = True
@_triton.jit
def _vecadd_kernel(x_ptr, y_ptr, out_ptr, n_elements, BLOCK: _tl.constexpr):
pid = _tl.program_id(axis=0)
offsets = pid * BLOCK + _tl.arange(0, BLOCK)
mask = offsets < n_elements
x = _tl.load(x_ptr + offsets, mask=mask, other=0)
y = _tl.load(y_ptr + offsets, mask=mask, other=0)
_tl.store(out_ptr + offsets, x + y, mask=mask)
except Exception:
_TRITON_AVAILABLE = False
_triton = None
_vecadd_kernel = None
def custom_kernel(data: input_t) -> output_t:
A, B, output = data
if not _TRITON_AVAILABLE:
output[...] = A + B
return output
if not (hasattr(A, "is_cuda") and A.is_cuda and hasattr(B, "is_cuda") and B.is_cuda):
output[...] = A + B
return output
if A.dtype != B.dtype or A.dtype != output.dtype:
output[...] = A + B
return output
if A.numel() != B.numel() or A.numel() != output.numel():
output[...] = A + B
return output
if not (A.is_contiguous() and B.is_contiguous() and output.is_contiguous()):
output[...] = A + B
return output
A_1d = A.view(-1)
B_1d = B.view(-1)
O_1d = output.view(-1)
n_elements = O_1d.numel()
grid = (_triton.cdiv(n_elements, 1024),)
_vecadd_kernel[grid](A_1d, B_1d, O_1d, n_elements, BLOCK=1024, num_warps=8)
return output
scrolls · 56 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON