Skip to content
KernelIndex
Search⌘K

submission 673457

Elán Zainos Corona · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 28 lines, June 9 Researcher Reciprocity License v1.0.

tuplas3.2.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-matmul-v2-673457?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp16

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
FP16 matmulsuite of 8 cases
NVIDIA B200
114.4µs
#15 of 53
2026-03-30

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:aacc8030804d49c52a6eb7720e4d4c94e052d857996e9a0f01c4b2ea0c062d0a
license declaredunknown
license concludedunknown
authorsElán Zainos Corona
imported2026-08-15

Kernel source

tuplas3.2.py28 lines
# -*- coding: utf-8 -*-
import os
# El buffer de 8MB es el requisito de la B200 para MatMul determinista
os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"

import torch

# Forzamos flags de alto rendimiento en el top-level
torch.backends.cuda.matmul.allow_tf32 = True

def custom_kernel(data):
    """
    Protocolo Tuplas v3.1 - Sentinel Omega.
    Indexación resiliente + Determinismo síncrono.
    """
    # Usar índices es más robusto que 'a, b, c = data'
    a = data[0]
    b = data[1]
    c = data[2]
    
    # Activamos determinismo justo antes del colapso (Cero Fricción)
    torch.use_deterministic_algorithms(True, warn_only=False)
    
    # La B200 utiliza Tensor Cores aquí de forma óptima
    torch.mm(a, b, out=c)
    
    return c

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Changes from previous submission

Against this author's previous submission submission 673439.

# -*- coding: utf-8 -*-
import os
+ # El buffer de 8MB es el requisito de la B200 para MatMul determinista
os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"
import torch
- # Optimizamos las banderas de hardware una sola vez a nivel global
+ # Forzamos flags de alto rendimiento en el top-level
torch.backends.cuda.matmul.allow_tf32 = True
- torch.backends.cudnn.deterministic = True
def custom_kernel(data):
"""
- MatMul Fractal - Motor de Alto Flujo.
+ Protocolo Tuplas v3.1 - Sentinel Omega.
+ Indexación resiliente + Determinismo síncrono.
"""
- # La trinidad (a, b, c) ocupa su posición en la memoria sin latencia de CPU.
- a, b, c = data
+ # Usar índices es más robusto que 'a, b, c = data'
+ a = data[0]
+ b = data[1]
+ c = data[2]
- # Recta de Vértices: Multiplicación in-place.
+ # Activamos determinismo justo antes del colapso (Cero Fricción)
+ torch.use_deterministic_algorithms(True, warn_only=False)
+
+ # La B200 utiliza Tensor Cores aquí de forma óptima
torch.mm(a, b, out=c)
return c
scrolls · 33 diff lines total

Best evidence level for this revision: reported

JSON