submission 180524
Nader · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 53 lines, June 9 Researcher Reciprocity License v1.0.
vectoradd_v2.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectoradd-v2-180524?include=source"interfacepython
Compatibility
measured onNVIDIA L4
declared hardwareNVIDIA L4
architecturessm_89
dtypesfp16
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:4137d69edc3fd54d44476d60d916c748bd7ec1fd9a7b3bcf2e6f7007832d442c
license declaredunknown
license concludedunknown
authorsNader
imported2026-08-15
Kernel source
vectoradd_v2.py53 lines
#!POPCORN leaderboard vectoradd_v2
# This is a submission template for popcorn leaderboard 'vectoradd_v2'.
# Your task is as follows:
# > Implement a float16 vector addition kernel.
# >
# > Input: tuple(torch.Tensor, torch.Tensor) with tensors of shape (N, N) and type torch.float16. These tensors are from
# > a normal distribution with mean 0 and variance 1.
# > Output: torch.Tensor of shape (N, N) and type torch.float16
# The deadline for this leaderboard is 2025-12-30 00:00:00+00:00
# You can automatically route this file to specific GPUs by adding a line
# `#!POPCORN gpus <GPUs>` to the header of this file.
# Happy hacking!
import subprocess
import torch
try:
import cuda.compute
except ImportError:
if "L4" in torch.cuda.get_device_name():
# Prefetch seems to be better on L4, so I'm using the potential-submissions branch
import os
subprocess.check_call(["rm", "-rf", "cccl"])
subprocess.check_call(["git", "clone", "--depth", "1", "--branch", "potential-submissions", "https://github.com/NaderAlAwar/cccl.git"])
subprocess.check_call(["git", "checkout", "83ad53b"], cwd="cccl")
env = os.environ.copy()
env["CC"] = "gcc"
env["CXX"] = "g++"
env["CMAKE_ARGS"] = "-DCMAKE_CXX_STANDARD=20 -DCMAKE_CUDA_STANDARD=20"
subprocess.check_call(["pip", "install", ".[cu12]"], cwd="cccl/python/cuda_cccl", env=env)
else:
subprocess.run(["pip", "install", "cuda-cccl[cu12]==0.4.3"])
from task import input_t, output_t
import cuda.compute
from cuda.compute import OpKind
build_A = torch.empty(2, 2, dtype=torch.float16, device="cuda")
build_B = torch.empty(2, 2, dtype=torch.float16, device="cuda")
build_output = torch.empty(2, 2, dtype=torch.float16, device="cuda")
transformer = cuda.compute.make_binary_transform(build_A, build_B, build_output, OpKind.PLUS)
def custom_kernel(data: input_t) -> output_t:
A, B, output = data
transformer(A, B, output, A.numel())
return output
scrolls · 53 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Changes from previous submission
Against this author's previous submission submission 180495.
Best evidence level for this revision: reported
JSON