Skip to content
KernelIndex
Search⌘K

submission 67265

shellsmile15795 · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 31 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectoradd-v2-67265?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp16

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
FP16 vector additionsuite of 5 cases
NVIDIA B200
237.0µs
#34 of 66
2025-11-06

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:27c6c9523289be4659561269543e44f7d56d9f341a280f6798b8d205b8691147
license declaredunknown
license concludedunknown
authorsshellsmile15795
imported2026-08-15

Kernel source

submission.py31 lines
from __future__ import annotations

from typing import Tuple

import torch

from task import input_t, output_t


def _validate_tensors(A: torch.Tensor, B: torch.Tensor, C: torch.Tensor) -> None:
    if not (A.is_cuda and B.is_cuda and C.is_cuda):
        raise RuntimeError("All tensors must reside on CUDA devices.")
    if A.dtype != torch.float16 or B.dtype != torch.float16 or C.dtype != torch.float16:
        raise RuntimeError("All tensors must use torch.float16 dtype.")
    if A.shape != B.shape or A.shape != C.shape:
        raise RuntimeError("Input and output tensors must share identical shapes.")
    if not (A.is_contiguous() and B.is_contiguous() and C.is_contiguous()):
        raise RuntimeError("Input and output tensors must be contiguous.")


def _launch_vector_add(A: torch.Tensor, B: torch.Tensor, C: torch.Tensor) -> None:
    # `torch.add` issues an efficient pointwise kernel and honors the provided output buffer.
    torch.add(A, B, out=C)


def custom_kernel(data: input_t) -> output_t:
    A, B, C = data  # type: ignore[misc]
    _validate_tensors(A, B, C)
    _launch_vector_add(A, B, C)
    return C
scrolls · 31 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON