Skip to content
KernelIndex
Search⌘K

submission 673503

Elán Zainos Corona · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 24 lines, June 9 Researcher Reciprocity License v1.0.

Par1.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-prefixsum-v2-673503?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
Inclusive prefix sumsuite of 11 cases
NVIDIA B200
2.20ms
#18 of 23
2026-03-30

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:3e74eb54ee1412526d19bce87170d725774a901a7d02d62fc0925703b002f10f
license declaredunknown
license concludedunknown
authorsElán Zainos Corona
imported2026-08-15

Kernel source

Par1.py24 lines
# -*- coding: utf-8 -*-
import os
os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"

import torch

# 1. PRE-CONFIGURACIÓN (World Line Stabilization)
# Seteamos esto una sola vez a nivel de módulo para ganar ~30µs por llamada.
torch.use_deterministic_algorithms(True)

def custom_kernel(data):
    """
    Sentinel Omega - Scan Inclusivo Turbo.
    """
    # 2. Recta de Vértices
    # Accedemos directamente a data[0]. No necesitamos el buffer data[1] 
    # si el protocolo exige un tensor de salida nuevo en float64.
    x = data[0]
    
    # 3. Colapso de Fase (Precision Control)
    # Cast y suma en una sola ráfaga de GPU.
    # El paso a float64 es ley para evitar el numerical instability (sqrt(n)).
    return torch.cumsum(x.to(torch.float64), dim=0)

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON