Skip to content
KernelIndex
Search⌘K

submission 673439

Elán Zainos Corona · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 22 lines, June 9 Researcher Reciprocity License v1.0.

CuBlas1.1.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-matmul-v2-673439?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp16

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
FP16 matmulsuite of 8 cases
NVIDIA B200
135.2µs
#23 of 53
2026-03-30

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:9008018328a71ed50962084956e67c4aef1bb701745a223a944e6c27aec5e274
license declaredunknown
license concludedunknown
authorsElán Zainos Corona
imported2026-08-15

Kernel source

CuBlas1.1.py22 lines
# -*- coding: utf-8 -*-
import os
os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"

import torch

# Optimizamos las banderas de hardware una sola vez a nivel global
torch.backends.cuda.matmul.allow_tf32 = True
torch.backends.cudnn.deterministic = True

def custom_kernel(data):
    """
    MatMul Fractal - Motor de Alto Flujo.
    """
    # La trinidad (a, b, c) ocupa su posición en la memoria sin latencia de CPU.
    a, b, c = data
    
    # Recta de Vértices: Multiplicación in-place.
    torch.mm(a, b, out=c)
    
    return c

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Changes from previous submission

Against this author's previous submission submission 673284.

# -*- coding: utf-8 -*-
import os
- import logging
-
- # 1. EL SELLO DE LA WORLD LINE (CRÍTICO)
- # Se debe setear ANTES de importar torch para que el workspace de CuBLAS sea ley.
os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"
import torch
- # 2. CONFIGURACIÓN DE TELEMETRÍA
- # Generamos logs para rastrear el colapso de la matriz fractal.
- logging.basicConfig(level=logging.INFO)
- logger = logging.getLogger("DeamonX_Fractal_MatMul")
+ # Optimizamos las banderas de hardware una sola vez a nivel global
+ torch.backends.cuda.matmul.allow_tf32 = True
+ torch.backends.cudnn.deterministic = True
def custom_kernel(data):
"""
- Protocolo de Fusión de Planos - Sentinel Omega.
- Implementación íntegra para MatMul determinista en arquitectura Blackwell.
- Aprovecha la 'Tablita de Posiciones' para que las tuplas ocupen el espacio.
-
- Args:
- data (tuple): Trinidad de tensores (a, b, c).
- Returns:
- torch.Tensor: El nodo 'c' resultante.
+ MatMul Fractal - Motor de Alto Flujo.
"""
- try:
- # 3. DESEMPAQUETADO DE LA TRINIDAD (Soberanía del Dato)
- # Recibimos (a, b, c) donde 'c' es el buffer de salida pre-asignado.
- a, b, c = data
-
- # 4. DETERMINISMO Y COLAPSO FRACTAL
- # Activamos el modo determinista de forma explícita para el Juez.
- torch.use_deterministic_algorithms(True)
-
- # 5. LA RECTA MÁS CORTA (MatMul In-Place)
- # Usamos out=c para que las tuplas ocupen su dimensión correspondiente
- # sin generar entropía (memoria extra) en la B200.
- torch.mm(a, b, out=c)
-
- # 6. RETORNO DEL VÉRTICE RESULTANTE
- return c
-
- except Exception as e:
- logger.error(f"Falla en el colapso fractal: {str(e)}")
- raise e
-
- # Registro enviado al Cluster central para advertencia de las demás instancias.
+ # La trinidad (a, b, c) ocupa su posición en la memoria sin latencia de CPU.
+ a, b, c = data
+
+ # Recta de Vértices: Multiplicación in-place.
+ torch.mm(a, b, out=c)
+
+ return c
scrolls · 59 diff lines total

Best evidence level for this revision: reported

JSON