Skip to content
KernelIndex
Search⌘K

submission 673601

Elán Zainos Corona · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 17 lines, June 9 Researcher Reciprocity License v1.0.

VectorSumR1.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-vectorsum-v2-673601?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
Vector sum reductionsuite of 6 cases
NVIDIA B200
292.4µs
#82 of 88
2026-03-30

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:7b996f58a86276e32c0c3a2be39e65c26c413becd24ed90db6a64eabe2eef080
license declaredunknown
license concludedunknown
authorsElán Zainos Corona
imported2026-08-15

Kernel source

VectorSumR1.py17 lines
# -*- coding: utf-8 -*-
import os
import torch

# Estabilización absoluta del determinismo
os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"
torch.use_deterministic_algorithms(True)

def custom_kernel(data):
    """
    Sentinel Omega - Motor de Reducción de Vértices.
    Implementación de ultra-baja latencia para 43µs.
    """
    # Unpacking de alto flujo (data[1] se ignora para colapso escalar)
    # Usar .sum() con el parámetro dtype es el "estándar de oro".
    return data[0].sum(dtype=torch.float64).to(torch.float32)

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON