Skip to content
KernelIndex
Search⌘K

submission 66401

CherryPie · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 61 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectoradd-v2-66401?include=source"
interfacepython
Compatibility
measured onNVIDIA H100
declared hardwareNVIDIA H100
architecturessm_90
dtypesfp16

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
FP16 vector additionsuite of 5 cases
NVIDIA H100
526.4µs
#23 of 44
2025-11-04

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:26fcd562a6a739e64d040ecd5c674c4fc3f7c45237214bf235641611af18f819
license declaredunknown
license concludedunknown
authorsCherryPie
imported2026-08-15

Kernel source

submission.py61 lines
#!POPCORN leaderboard vectoradd_v2
import torch
import triton
import triton.language as tl


@triton.jit
def add_kernel(
    a_ptr,  # pointer to input A
    b_ptr,  # pointer to input B
    output_ptr,  # pointer to output
    n_elements,  # total number of elements
    BLOCK_SIZE: tl.constexpr,  # number of elements per block
):
    """Optimized element-wise addition kernel using Triton"""
    # Get program ID and compute element offset
    pid = tl.program_id(axis=0)
    block_start = pid * BLOCK_SIZE
    offsets = block_start + tl.arange(0, BLOCK_SIZE)
    
    # Create mask for boundary checking
    mask = offsets < n_elements
    
    # Load data with masking
    a = tl.load(a_ptr + offsets, mask=mask)
    b = tl.load(b_ptr + offsets, mask=mask)
    
    # Perform addition
    output = a + b
    
    # Store result with masking
    tl.store(output_ptr + offsets, output, mask=mask)


def custom_kernel(data):
    """
    High-performance vector addition using Triton CUDA kernel.
    Args:
        data: Tuple of tensors (A, B, output)
    Returns:
        output tensor with A + B
    """
    A, B, output = data
    
    # Get total number of elements
    n_elements = output.numel()
    
    # Choose optimal block size (tuned for modern GPUs)
    BLOCK_SIZE = 1024
    
    # Calculate grid size
    grid = lambda meta: (triton.cdiv(n_elements, meta['BLOCK_SIZE']),)
    
    # Launch kernel
    add_kernel[grid](
        A, B, output,
        n_elements,
        BLOCK_SIZE=BLOCK_SIZE,
    )
    
    return output
scrolls · 61 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON