Skip to content
KernelIndex
Search⌘K

submission 779877

ajay_a · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 48 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-matmul-v2-779877?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp16

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
FP16 matmulsuite of 8 cases
NVIDIA B200
115.3µs
#18 of 53
2026-04-23

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:71ab6a72e77eb903c451c0eff590266e743bf0d020ab14bdfcd7234b7f88d8a7
license declaredunknown
license concludedunknown
authorsajay_a
imported2026-08-15

Kernel source

submission.py48 lines
#!POPCORN leaderboard matmul_v2
#!POPCORN gpu B200

# at::matmul C++ wrapper. Same winning pattern as conv2d_v2:
# - Dispatch to cuBLAS via ATen (same function the bot's reference uses)
# - Force TF32 off + deterministic to match reference precision exactly
# - cudnn.benchmark=True so the cuBLASlt heuristic can pick the best algo
from task import input_t, output_t
import torch

torch.backends.cudnn.allow_tf32 = False
torch.backends.cudnn.deterministic = True
torch.backends.cudnn.benchmark = True
torch.backends.cuda.matmul.allow_tf32 = False

from torch.utils.cpp_extension import load_inline


_CUDA_SRC = r"""
#include <ATen/ATen.h>
#include <torch/torch.h>

// Dispatches to cuBLAS for dense matmul; write result into preallocated out.
void matmul_fwd(const torch::Tensor& A,
                const torch::Tensor& B,
                torch::Tensor& out) {
    at::matmul_out(out, A, B);
}
"""

_CPP_SRC = "void matmul_fwd(const torch::Tensor&, const torch::Tensor&, torch::Tensor&);"

_mod = load_inline(
    name="matmul_cublas_wrap",
    cpp_sources=_CPP_SRC,
    cuda_sources=_CUDA_SRC,
    functions=["matmul_fwd"],
    extra_cuda_cflags=["-O3", "-arch=sm_100"],
    extra_cflags=["-O3"],
    verbose=False,
)


def custom_kernel(data: input_t) -> output_t:
    A, B, out = data
    _mod.matmul_fwd(A, B, out)
    return out
scrolls · 48 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Changes from previous submission

Against this author's previous submission submission 779854.

Best evidence level for this revision: reported

JSON