submission 512763
mreso · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 45 lines, June 9 Researcher Reciprocity License v1.0.
submission_conv2d_v2.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-conv2d-v2-512763?include=source"interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp32
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:38b2a0397745faeb1e1416fe2825fa3e3714beeaae45f6e17e60dd42e5dca8db
license declaredunknown
license concludedunknown
authorsmreso
imported2026-08-15
Kernel source
submission_conv2d_v2.py45 lines
# submission_conv2d_v2.py
# 2D convolution: out = conv2d(inp, kernel, stride=1, padding=0)
# Interface: custom_kernel((inp, kernel, out)) -> out
# inp: (B, C_in, H, W) float32
# kernel: (C_out, C_in, kH, kW) float32
# out: (B, C_out, H-kH+1, W-kW+1) float32
#
# Strategy: im2col (via F.unfold) + batched GEMM (torch.bmm / cuBLAS).
# A pure-Triton float32 GEMM using tl.dot operates in tf32 on Ampere+ and
# produces errors > rtol/atol=1e-3 for this problem. torch.bmm uses cuBLAS
# full float32 (or configurable tf32) and stays within tolerance.
#
# im2col: F.unfold(inp, (kH, kW)) → (B, K, H_out*W_out) where K=C_in*kH*kW
# GEMM: kernel.view(C_out, K) @ unfold → (B, C_out, H_out*W_out)
# Reshape: → (B, C_out, H_out, W_out)
import torch
import torch.nn.functional as F
from task import input_t, output_t
def custom_kernel(data: input_t) -> output_t:
inp, kernel, out = data
B, C_in, H, W = inp.shape
C_out, _, kH, kW = kernel.shape
K = C_in * kH * kW
H_out = H - kH + 1
W_out = W - kW + 1
# ── im2col ───────────────────────────────────────────────────────────────
# cols: (B, K, H_out*W_out)
cols = F.unfold(inp, (kH, kW))
# ── batched GEMM ─────────────────────────────────────────────────────────
# (B, C_out, K) @ (B, K, H_out*W_out) → (B, C_out, H_out*W_out)
k_mat = kernel.reshape(C_out, K)
result = torch.bmm(
k_mat.unsqueeze(0).expand(B, -1, -1), # (B, C_out, K)
cols, # (B, K, H_out*W_out)
)
out.copy_(result.reshape(B, C_out, H_out, W_out))
return out
scrolls · 45 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON