submission 779877
ajay_a · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 48 lines, June 9 Researcher Reciprocity License v1.0.
submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-matmul-v2-779877?include=source"interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp16
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:71ab6a72e77eb903c451c0eff590266e743bf0d020ab14bdfcd7234b7f88d8a7
license declaredunknown
license concludedunknown
authorsajay_a
imported2026-08-15
Kernel source
submission.py48 lines
#!POPCORN leaderboard matmul_v2
#!POPCORN gpu B200
# at::matmul C++ wrapper. Same winning pattern as conv2d_v2:
# - Dispatch to cuBLAS via ATen (same function the bot's reference uses)
# - Force TF32 off + deterministic to match reference precision exactly
# - cudnn.benchmark=True so the cuBLASlt heuristic can pick the best algo
from task import input_t, output_t
import torch
torch.backends.cudnn.allow_tf32 = False
torch.backends.cudnn.deterministic = True
torch.backends.cudnn.benchmark = True
torch.backends.cuda.matmul.allow_tf32 = False
from torch.utils.cpp_extension import load_inline
_CUDA_SRC = r"""
#include <ATen/ATen.h>
#include <torch/torch.h>
// Dispatches to cuBLAS for dense matmul; write result into preallocated out.
void matmul_fwd(const torch::Tensor& A,
const torch::Tensor& B,
torch::Tensor& out) {
at::matmul_out(out, A, B);
}
"""
_CPP_SRC = "void matmul_fwd(const torch::Tensor&, const torch::Tensor&, torch::Tensor&);"
_mod = load_inline(
name="matmul_cublas_wrap",
cpp_sources=_CPP_SRC,
cuda_sources=_CUDA_SRC,
functions=["matmul_fwd"],
extra_cuda_cflags=["-O3", "-arch=sm_100"],
extra_cflags=["-O3"],
verbose=False,
)
def custom_kernel(data: input_t) -> output_t:
A, B, out = data
_mod.matmul_fwd(A, B, out)
return out
scrolls · 48 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Changes from previous submission
Against this author's previous submission submission 779854.
Best evidence level for this revision: reported
JSON