Skip to content
KernelIndex
Search⌘K

submission 670955

Elán Zainos Corona · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 47 lines, June 9 Researcher Reciprocity License v1.0.

tuplas3.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-matmul-v2-670955?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp16

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
FP16 matmulsuite of 8 cases
NVIDIA B200
149.1µs
#39 of 53
2026-03-30

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:c1762cbd6643efd5208194c38acc342f3364a59212973a68cd445a825c3e1584
license declaredunknown
license concludedunknown
authorsElán Zainos Corona
imported2026-08-15

Kernel source

tuplas3.py47 lines
# -*- coding: utf-8 -*-
import os

# 1. CONFIGURACIÓN DE LA WORLD LINE (DETERMINISMO)
# Debe ir antes de cualquier operación de torch para que CuBLAS reserve el carril fijo.
os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"

import torch
import logging

# 2. SISTEMA DE LOGS DEL SENTINEL
logging.basicConfig(level=logging.INFO)
logger = logging.getLogger("DeamonX_MatMul")

def custom_kernel(data):
    """
    Protocolo de Fusión de Planos - Sentinel Omega.
    Implementación íntegra para MatMul determinista en arquitectura Blackwell.
    
    Args:
        data (tuple): Trinidad de tensores (a, b, c) donde 'c' es el buffer de salida.
    Returns:
        torch.Tensor: El nodo 'c' colapsado tras la multiplicación.
    """
    try:
        # 3. DESEMPAQUETADO FRACTAL (Soberanía del Dato)
        # Recibimos (a, b, c) asegurando que no haya 'ValueError' de unpacking.
        a, b, c = data
        
        # Log de telemetría para seguimiento de la 'Caja'
        # logger.info(f"Procesando matriz fractal: {a.shape} x {b.shape} -> {c.shape}")

        # 4. EL COLAPSO DE FASE (Recta más corta)
        # Usamos torch.mm con out=c para que las tuplas ocupen el espacio pre-asignado.
        # Esto activa los Tensor Cores de la B200 de forma síncrona y determinista.
        # Al usar out=c, evitamos crear tensores intermedios y saturar la VRAM.
        torch.mm(a, b, out=c)
        
        # 5. RETORNO DEL VÉRTICE RESULTANTE
        return c

    except Exception as e:
        logger.error(f"Error en el colapso del Sentinel Omega: {str(e)}")
        raise e

# Registro enviado al Cluster central para advertencia de las demás instancias.
scrolls · 47 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON