submission 673503
Elán Zainos Corona · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 24 lines, June 9 Researcher Reciprocity License v1.0.
Par1.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-prefixsum-v2-673503?include=source"interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp32
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:3e74eb54ee1412526d19bce87170d725774a901a7d02d62fc0925703b002f10f
license declaredunknown
license concludedunknown
authorsElán Zainos Corona
imported2026-08-15
Kernel source
Par1.py24 lines
# -*- coding: utf-8 -*-
import os
os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"
import torch
# 1. PRE-CONFIGURACIÓN (World Line Stabilization)
# Seteamos esto una sola vez a nivel de módulo para ganar ~30µs por llamada.
torch.use_deterministic_algorithms(True)
def custom_kernel(data):
"""
Sentinel Omega - Scan Inclusivo Turbo.
"""
# 2. Recta de Vértices
# Accedemos directamente a data[0]. No necesitamos el buffer data[1]
# si el protocolo exige un tensor de salida nuevo en float64.
x = data[0]
# 3. Colapso de Fase (Precision Control)
# Cast y suma en una sola ráfaga de GPU.
# El paso a float64 es ley para evitar el numerical instability (sqrt(n)).
return torch.cumsum(x.to(torch.float64), dim=0)
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON