Skip to content
KernelIndex
Search⌘K

submission 674172

Elán Zainos Corona · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 31 lines, June 9 Researcher Reciprocity License v1.0.

tuplas3.4.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-matmul-v2-674172?include=source"
interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp16

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
FP16 matmulsuite of 8 cases
NVIDIA A100
641.5µs
#5 of 27
2026-03-30

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:0f1cb02aa30fdae4fbbe32dde75ff923a9a630e6c1791ddc5ac535d9b13921c4
license declaredunknown
license concludedunknown
authorsElán Zainos Corona
imported2026-08-15

Kernel source

tuplas3.4.py31 lines
# -*- coding: utf-8 -*-
import os
import logging
import torch

os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"

logging.basicConfig(level=logging.INFO)
logger = logging.getLogger("DeamonX_Tuplas_B200")

def custom_kernel(data):
    """
    Protocolo Tuplas - Corrección de alineación de memoria.
    """
    try:
        # Forzar alineación de memoria contigua para optimizar el bus de datos
        a = data[0].contiguous()
        b = data[1].contiguous()
        c = data[2]
        
        torch.use_deterministic_algorithms(True, warn_only=False)
        
        # Ejecución in-place
        torch.mm(a, b, out=c)
        
        return c
        
    except Exception as e:
        logger.error(f"Falla de ejecución en Tuplas: {str(e)}")
        raise e
scrolls · 31 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Changes from previous submission

Against this author's previous submission submission 673457.

# -*- coding: utf-8 -*-
import os
- # El buffer de 8MB es el requisito de la B200 para MatMul determinista
- os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"
-
+ import logging
import torch
- # Forzamos flags de alto rendimiento en el top-level
- torch.backends.cuda.matmul.allow_tf32 = True
+ os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"
+ logging.basicConfig(level=logging.INFO)
+ logger = logging.getLogger("DeamonX_Tuplas_B200")
+
def custom_kernel(data):
"""
- Protocolo Tuplas v3.1 - Sentinel Omega.
- Indexación resiliente + Determinismo síncrono.
+ Protocolo Tuplas - Corrección de alineación de memoria.
"""
- # Usar índices es más robusto que 'a, b, c = data'
- a = data[0]
- b = data[1]
- c = data[2]
-
- # Activamos determinismo justo antes del colapso (Cero Fricción)
- torch.use_deterministic_algorithms(True, warn_only=False)
-
- # La B200 utiliza Tensor Cores aquí de forma óptima
- torch.mm(a, b, out=c)
-
- return c
+ try:
+ # Forzar alineación de memoria contigua para optimizar el bus de datos
+ a = data[0].contiguous()
+ b = data[1].contiguous()
+ c = data[2]
+
+ torch.use_deterministic_algorithms(True, warn_only=False)
+
+ # Ejecución in-place
+ torch.mm(a, b, out=c)
+
+ return c
+
+ except Exception as e:
+ logger.error(f"Falla de ejecución en Tuplas: {str(e)}")
+ raise e
scrolls · 49 diff lines total

Best evidence level for this revision: reported

JSON