Skip to content
KernelIndex
Search⌘K

submission 512763

mreso · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 45 lines, June 9 Researcher Reciprocity License v1.0.

submission_conv2d_v2.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-conv2d-v2-512763?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
2D convolutionsuite of 5 cases
NVIDIA B200
42.1ms
#15 of 28
2026-03-04

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:38b2a0397745faeb1e1416fe2825fa3e3714beeaae45f6e17e60dd42e5dca8db
license declaredunknown
license concludedunknown
authorsmreso
imported2026-08-15

Kernel source

submission_conv2d_v2.py45 lines
# submission_conv2d_v2.py
# 2D convolution: out = conv2d(inp, kernel, stride=1, padding=0)
# Interface: custom_kernel((inp, kernel, out)) -> out
#   inp:    (B, C_in,  H,    W)    float32
#   kernel: (C_out, C_in, kH, kW)  float32
#   out:    (B, C_out, H-kH+1, W-kW+1) float32
#
# Strategy: im2col (via F.unfold) + batched GEMM (torch.bmm / cuBLAS).
# A pure-Triton float32 GEMM using tl.dot operates in tf32 on Ampere+ and
# produces errors > rtol/atol=1e-3 for this problem.  torch.bmm uses cuBLAS
# full float32 (or configurable tf32) and stays within tolerance.
#
# im2col: F.unfold(inp, (kH, kW)) → (B, K, H_out*W_out)  where K=C_in*kH*kW
# GEMM:   kernel.view(C_out, K)   @ unfold → (B, C_out, H_out*W_out)
# Reshape: → (B, C_out, H_out, W_out)

import torch
import torch.nn.functional as F
from task import input_t, output_t


def custom_kernel(data: input_t) -> output_t:
    inp, kernel, out = data

    B, C_in, H, W   = inp.shape
    C_out, _, kH, kW = kernel.shape
    K     = C_in * kH * kW
    H_out = H - kH + 1
    W_out = W - kW + 1

    # ── im2col ───────────────────────────────────────────────────────────────
    # cols: (B, K, H_out*W_out)
    cols = F.unfold(inp, (kH, kW))

    # ── batched GEMM ─────────────────────────────────────────────────────────
    # (B, C_out, K) @ (B, K, H_out*W_out) → (B, C_out, H_out*W_out)
    k_mat  = kernel.reshape(C_out, K)
    result = torch.bmm(
        k_mat.unsqueeze(0).expand(B, -1, -1),  # (B, C_out, K)
        cols,                                   # (B, K, H_out*W_out)
    )

    out.copy_(result.reshape(B, C_out, H_out, W_out))
    return out
scrolls · 45 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON