submission 669289
Gurkirat Singh · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 37 lines, June 9 Researcher Reciprocity License v1.0.
submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectoradd-v2-669289?include=source"interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp16
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:101bf10768812371b8fd2e172a301e44cbdd59c80e91d7c0208b2fe1e449e7dd
license declaredunknown
license concludedunknown
authorsGurkirat Singh
imported2026-08-15
Kernel source
submission.py37 lines
# Write your code here#!POPCORN leaderboard vectoradd_v2
# This is a submission template for popcorn leaderboard 'vectoradd'.
# Your task is as follows:
# > Implement a float16 vector addition kernel.
# >
# > Input: tuple(torch.Tensor, torch.Tensor) with tensors of shape (N, N) and type torch.float16. These tensors are from
# > a normal distribution with mean 0 and variance 1.
# > Output: torch.Tensor of shape (N, N) and type torch.float16
# The deadline for this leaderboard is 2025-04-30 00:00:00+00:00
# You can automatically route this file to specific GPUs by adding a line
# `#!POPCORN gpus <GPUs>` to the header of this file.
# Happy hacking!
from task import input_t, output_t
def custom_kernel(data: input_t) -> output_t:
"""
Custom kernel for vector addition.
"""
# Unpack the input tuple
s = data[0]
t = data[1]
# Perform element-wise addition
q = s + t
# Return the result
return q
scrolls · 37 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Changes from previous submission
Against this author's previous submission submission 669062.
+ # Write your code here#!POPCORN leaderboard vectoradd_v2+ # This is a submission template for popcorn leaderboard 'vectoradd'.+ # Your task is as follows:+ # > Implement a float16 vector addition kernel.- import torch- import triton- import triton.language as tl- from task import input_t, output_t+ # >+ # > Input: tuple(torch.Tensor, torch.Tensor) with tensors of shape (N, N) and type torch.float16. These tensors are from- @triton.jit- def add_kernel(- A_ptr, B_ptr, C_ptr, M, N,- BLOCK_SIZE: tl.constexpr,- ):- row_idx = tl.program_id(0) * BLOCK_SIZE + tl.arange(0, BLOCK_SIZE)- col_idx = tl.program_id(1) * BLOCK_SIZE + tl.arange(0, BLOCK_SIZE)-- mask_row = row_idx < M- mask_col = col_idx < N+ # > a normal distribution with mean 0 and variance 1.- A = tl.load(A_ptr + row_idx[:, None] * N + col_idx[None, :], mask=mask_row[:, None] & mask_col[None, :], other=0.0)- B = tl.load(B_ptr + row_idx[:, None] * N + col_idx[None, :], mask=mask_row[:, None] & mask_col[None, :], other=0.0)+ # > Output: torch.Tensor of shape (N, N) and type torch.float16- C = A + B- tl.store(C_ptr + row_idx[:, None] * N + col_idx[None, :], C, mask=mask_row[:, None] & mask_col[None, :])+ # The deadline for this leaderboard is 2025-04-30 00:00:00+00:00- def custom_kernel(data: input_t) -> output_t:- A, B, C = data- M, N = A.shape+ # You can automatically route this file to specific GPUs by adding a line+ # `#!POPCORN gpus <GPUs>` to the header of this file.+ # Happy hacking!- BLOCK_SIZE = 32- grid = (triton.cdiv(M, BLOCK_SIZE), triton.cdiv(N, BLOCK_SIZE))+ from task import input_t, output_t- add_kernel[grid](- A, B, C, M, N,- BLOCK_SIZE=BLOCK_SIZE,- )- return CNo newline at end of file+ def custom_kernel(data: input_t) -> output_t:+ """+ Custom kernel for vector addition.+ """+ # Unpack the input tuple+ s = data[0]+ t = data[1]++ # Perform element-wise addition+ q = s + t++ # Return the result+ return q
scrolls · 65 diff lines total
Best evidence level for this revision: reported
JSON