Skip to content
KernelIndex
Search⌘K

submission 669289

Gurkirat Singh · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 37 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectoradd-v2-669289?include=source"
interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp16

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
FP16 vector additionsuite of 5 cases
NVIDIA A100
894.3µs
#5= of 87
2026-03-30

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:101bf10768812371b8fd2e172a301e44cbdd59c80e91d7c0208b2fe1e449e7dd
license declaredunknown
license concludedunknown
authorsGurkirat Singh
imported2026-08-15

Kernel source

submission.py37 lines
# Write your code here#!POPCORN leaderboard vectoradd_v2

# This is a submission template for popcorn leaderboard 'vectoradd'.
# Your task is as follows:
# > Implement a float16 vector addition kernel.

# >

# > Input: tuple(torch.Tensor, torch.Tensor) with tensors of shape (N, N) and type torch.float16. These tensors are from

# > a normal distribution with mean 0 and variance 1.

# > Output: torch.Tensor of shape (N, N) and type torch.float16

# The deadline for this leaderboard is 2025-04-30 00:00:00+00:00

# You can automatically route this file to specific GPUs by adding a line
# `#!POPCORN gpus <GPUs>` to the header of this file.
# Happy hacking!

from task import input_t, output_t


def custom_kernel(data: input_t) -> output_t:
    """
    Custom kernel for vector addition.
    """
    # Unpack the input tuple
    s = data[0]
    t = data[1]

    # Perform element-wise addition
    q = s + t

    # Return the result
    return q
scrolls · 37 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Changes from previous submission

Against this author's previous submission submission 669062.

+ # Write your code here#!POPCORN leaderboard vectoradd_v2
+ # This is a submission template for popcorn leaderboard 'vectoradd'.
+ # Your task is as follows:
+ # > Implement a float16 vector addition kernel.
- import torch
- import triton
- import triton.language as tl
- from task import input_t, output_t
+ # >
+ # > Input: tuple(torch.Tensor, torch.Tensor) with tensors of shape (N, N) and type torch.float16. These tensors are from
- @triton.jit
- def add_kernel(
- A_ptr, B_ptr, C_ptr, M, N,
- BLOCK_SIZE: tl.constexpr,
- ):
- row_idx = tl.program_id(0) * BLOCK_SIZE + tl.arange(0, BLOCK_SIZE)
- col_idx = tl.program_id(1) * BLOCK_SIZE + tl.arange(0, BLOCK_SIZE)
-
- mask_row = row_idx < M
- mask_col = col_idx < N
+ # > a normal distribution with mean 0 and variance 1.
- A = tl.load(A_ptr + row_idx[:, None] * N + col_idx[None, :], mask=mask_row[:, None] & mask_col[None, :], other=0.0)
- B = tl.load(B_ptr + row_idx[:, None] * N + col_idx[None, :], mask=mask_row[:, None] & mask_col[None, :], other=0.0)
+ # > Output: torch.Tensor of shape (N, N) and type torch.float16
- C = A + B
- tl.store(C_ptr + row_idx[:, None] * N + col_idx[None, :], C, mask=mask_row[:, None] & mask_col[None, :])
+ # The deadline for this leaderboard is 2025-04-30 00:00:00+00:00
- def custom_kernel(data: input_t) -> output_t:
- A, B, C = data
- M, N = A.shape
+ # You can automatically route this file to specific GPUs by adding a line
+ # `#!POPCORN gpus <GPUs>` to the header of this file.
+ # Happy hacking!
- BLOCK_SIZE = 32
- grid = (triton.cdiv(M, BLOCK_SIZE), triton.cdiv(N, BLOCK_SIZE))
+ from task import input_t, output_t
- add_kernel[grid](
- A, B, C, M, N,
- BLOCK_SIZE=BLOCK_SIZE,
- )
- return C
No newline at end of file
+ def custom_kernel(data: input_t) -> output_t:
+ """
+ Custom kernel for vector addition.
+ """
+ # Unpack the input tuple
+ s = data[0]
+ t = data[1]
+
+ # Perform element-wise addition
+ q = s + t
+
+ # Return the result
+ return q
scrolls · 65 diff lines total

Best evidence level for this revision: reported

JSON