submission 191358
Nader · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 65 lines, June 9 Researcher Reciprocity License v1.0.
vectoradd_v2.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectoradd-v2-191358?include=source"interfacepython
Compatibility
measured onNVIDIA H100
declared hardwareNVIDIA H100
architecturessm_90
dtypesfp16
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:491298ead10121a95f6a67c02c90c4e3ce4186a2013ec50c55e1fbd0d2a8ae9c
license declaredunknown
license concludedunknown
authorsNader
imported2026-08-15
Kernel source
vectoradd_v2.py65 lines
#!POPCORN leaderboard vectoradd_v2
# This is a submission template for popcorn leaderboard 'vectoradd_v2'.
# Your task is as follows:
# > Implement a float16 vector addition kernel.
# >
# > Input: tuple(torch.Tensor, torch.Tensor) with tensors of shape (N, N) and type torch.float16. These tensors are from
# > a normal distribution with mean 0 and variance 1.
# > Output: torch.Tensor of shape (N, N) and type torch.float16
# The deadline for this leaderboard is 2025-12-30 00:00:00+00:00
# You can automatically route this file to specific GPUs by adding a line
# `#!POPCORN gpus <GPUs>` to the header of this file.
# Happy hacking!
import subprocess
import torch
try:
import cuda.compute
except ImportError:
if "L4" in torch.cuda.get_device_name():
# Prefetch seems to be better on L4, so I'm using the potential-submissions branch
import os
subprocess.check_call(["rm", "-rf", "cccl"])
subprocess.check_call(["git", "clone", "--depth", "1", "--branch", "potential-submissions", "https://github.com/NaderAlAwar/cccl.git"])
subprocess.check_call(["git", "checkout", "83ad53b"], cwd="cccl")
env = os.environ.copy()
env["CC"] = "gcc"
env["CXX"] = "g++"
env["CMAKE_ARGS"] = "-DCMAKE_CXX_STANDARD=20 -DCMAKE_CUDA_STANDARD=20"
subprocess.check_call(["pip", "install", ".[cu12]"], cwd="cccl/python/cuda_cccl", env=env)
elif "H100" in torch.cuda.get_device_name() or "B200" in torch.cuda.get_device_name() or "A100" in torch.cuda.get_device_name():
import os
# subprocess.check_call(["rm", "-rf", "cccl"])
subprocess.run(["git", "clone", "--depth", "2", "--branch", "potential-submissions", "https://github.com/NaderAlAwar/cccl.git"])
subprocess.check_call(["git", "checkout", "e063657650987d055d1765a81395b0a4f2302b4e"], cwd="cccl")
env = os.environ.copy()
env["CC"] = "gcc"
env["CXX"] = "g++"
env["CMAKE_ARGS"] = "-DCMAKE_CXX_STANDARD=20 -DCMAKE_CUDA_STANDARD=20"
subprocess.check_call(["pip", "install", "-e", ".[cu12]"], cwd="cccl/python/cuda_cccl", env=env)
else:
subprocess.run(["pip", "install", "cuda-cccl[cu12]==0.4.3"])
from task import input_t, output_t
import cuda.compute
from cuda.compute import OpKind
build_A = torch.empty(2, 2, dtype=torch.float16, device="cuda")
build_B = torch.empty(2, 2, dtype=torch.float16, device="cuda")
build_output = torch.empty(2, 2, dtype=torch.float16, device="cuda")
transformer = cuda.compute.make_binary_transform(build_A, build_B, build_output, OpKind.PLUS)
def custom_kernel(data: input_t) -> output_t:
A, B, output = data
transformer(A, B, output, A.numel())
return output
scrolls · 65 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Changes from previous submission
Against this author's previous submission submission 191357.
Best evidence level for this revision: reported
JSON