Skip to content
KernelIndex
Search⌘K

submission 191385

Nader · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 65 lines, June 9 Researcher Reciprocity License v1.0.

vectoradd_v2.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectoradd-v2-191385?include=source"
interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp16

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
FP16 vector additionsuite of 5 cases
NVIDIA A100
892.9µs
#2 of 87
2025-12-22

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:fc04748c7ae90039cffedf6752550493cc520a2556176c0384e7ffb4bac78286
license declaredunknown
license concludedunknown
authorsNader
imported2026-08-15

Kernel source

vectoradd_v2.py65 lines
#!POPCORN leaderboard vectoradd_v2

# This is a submission template for popcorn leaderboard 'vectoradd_v2'.
# Your task is as follows:
# > Implement a float16 vector addition kernel.
# > 
# > Input: tuple(torch.Tensor, torch.Tensor) with tensors of shape (N, N) and type torch.float16. These tensors are from
# > a normal distribution with mean 0 and variance 1.
# > Output: torch.Tensor of shape (N, N) and type torch.float16
# The deadline for this leaderboard is 2025-12-30 00:00:00+00:00

# You can automatically route this file to specific GPUs by adding a line
# `#!POPCORN gpus <GPUs>` to the header of this file.
# Happy hacking!

import subprocess
import torch

try:
    import cuda.compute
except ImportError:
    if "L4" in torch.cuda.get_device_name():
        # Prefetch seems to be better on L4, so I'm using the potential-submissions branch
        import os
        subprocess.check_call(["rm", "-rf", "cccl"])
        subprocess.check_call(["git", "clone", "--depth", "1", "--branch", "potential-submissions", "https://github.com/NaderAlAwar/cccl.git"])
        subprocess.check_call(["git", "checkout", "83ad53b"], cwd="cccl")

        env = os.environ.copy()
        env["CC"] = "gcc"
        env["CXX"] = "g++"
        env["CMAKE_ARGS"] = "-DCMAKE_CXX_STANDARD=20 -DCMAKE_CUDA_STANDARD=20"

        subprocess.check_call(["pip", "install", ".[cu12]"], cwd="cccl/python/cuda_cccl", env=env)
    elif "H100" in torch.cuda.get_device_name() or "B200" in torch.cuda.get_device_name() or "A100" in torch.cuda.get_device_name():
        import os
        # subprocess.check_call(["rm", "-rf", "cccl"])
        subprocess.run(["git", "clone", "--depth", "2", "--branch", "potential-submissions", "https://github.com/NaderAlAwar/cccl.git"])
        subprocess.check_call(["git", "checkout", "e063657650987d055d1765a81395b0a4f2302b4e"], cwd="cccl")

        env = os.environ.copy()
        env["CC"] = "gcc"
        env["CXX"] = "g++"
        env["CMAKE_ARGS"] = "-DCMAKE_CXX_STANDARD=20 -DCMAKE_CUDA_STANDARD=20"

        subprocess.check_call(["pip", "install", "-e", ".[cu12]"], cwd="cccl/python/cuda_cccl", env=env)
    else:
        subprocess.run(["pip", "install", "cuda-cccl[cu12]==0.4.3"])

from task import input_t, output_t

import cuda.compute
from cuda.compute import OpKind

build_A = torch.empty(2, 2, dtype=torch.float16, device="cuda")
build_B = torch.empty(2, 2, dtype=torch.float16, device="cuda")
build_output = torch.empty(2, 2, dtype=torch.float16, device="cuda")

transformer = cuda.compute.make_binary_transform(build_A, build_B, build_output, OpKind.PLUS)

def custom_kernel(data: input_t) -> output_t:
    A, B, output = data
    transformer(A, B, output, A.numel())
    return output
scrolls · 65 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Changes from previous submission

Against this author's previous submission submission 191376.

Best evidence level for this revision: reported

JSON