Skip to content
KernelIndex
Search⌘K

submission 66194

Vishal369 · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 77 lines, June 9 Researcher Reciprocity License v1.0.

sssss.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectoradd-v2-66194?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp16

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
FP16 vector additionsuite of 5 cases
NVIDIA B200
238.2µs
#40 of 66
2025-10-17

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:9af4b0af00d0a14befe2e19b736da35c1a7b85c9a6b5f67a2a068ab1ba7485cd
license declaredunknown
license concludedunknown
authorsVishal369
imported2026-08-15

Kernel source

sssss.py77 lines
import torch
import triton
import triton.language as tl
 
@triton.jit
def vectoradd_kernel(
    a_ptr,  # Pointer to first input tensor
    b_ptr,  # Pointer to second input tensor
    c_ptr,  # Pointer to output tensor
    N,      # Total number of elements
    BLOCK_SIZE: tl.constexpr,  # Number of elements per block
):
    """
    Optimized float16 vector addition kernel.
    Each program processes BLOCK_SIZE elements.
    """
    # Get program ID
    pid = tl.program_id(0)
   
    # Compute block start offset
    block_start = pid * BLOCK_SIZE
   
    # Generate offsets for this block
    offsets = block_start + tl.arange(0, BLOCK_SIZE)
   
    # Create mask for boundary checking (handles non-multiple of BLOCK_SIZE)
    mask = offsets < N
   
    # Load data with vectorized operations (float16)
    # Use eviction policy for better cache utilization
    a = tl.load(a_ptr + offsets, mask=mask, other=0.0, eviction_policy='evict_first')
    b = tl.load(b_ptr + offsets, mask=mask, other=0.0, eviction_policy='evict_first')
   
    # Perform addition
    c = a + b
   
    # Store result with vectorized write
    tl.store(c_ptr + offsets, c, mask=mask)
 
 
def custom_kernel(data):
    """
    Main kernel function that launches the Triton kernel.
   
    Args:
        data: Tuple of (A, B, C) where A and B are input tensors,
              and C is the pre-allocated output tensor.
   
    Returns:
        C: Output tensor containing element-wise sum
    """
    A, B, C = data
   
    # Ensure inputs are contiguous for optimal memory access
    A = A.contiguous()
    B = B.contiguous()
   
    # Get tensor dimensions
    N = A.numel()
   
    # Choose optimal block size based on GPU architecture
    # Larger block sizes for better memory coalescing and fewer kernel launches
    # Powers of 2 work best for GPU memory systems
    BLOCK_SIZE = 1024  # Optimal for most modern GPUs (A100, H100, B200)
   
    # Calculate grid size (number of programs to launch)
    grid = lambda meta: (triton.cdiv(N, meta['BLOCK_SIZE']),)
   
    # Launch kernel
    vectoradd_kernel[grid](
        A, B, C, N,
        BLOCK_SIZE=BLOCK_SIZE,
    )
   
    return C
 
 
scrolls · 77 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON