submission 512856
JordanNanos · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 25 lines, June 9 Researcher Reciprocity License v1.0.
conv2d_py_submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-conv2d-v2-512856?include=source"interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp32
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:4b337fcdf73f8b9487b2e7dd0c2e083bb7878c2c41a366cbe1a8a692cdc7cae9
license declaredunknown
license concludedunknown
authorsJordanNanos
imported2026-08-15
Kernel source
conv2d_py_submission.py25 lines
import torch
import torch.nn.functional as F
from task import input_t, output_t
def custom_kernel(data: input_t) -> output_t:
input_tensor, kernel, output = data
batch, channels, height, width = input_tensor.shape
out_channels, in_channels, kh, kw = kernel.shape
out_h = height - kh + 1
out_w = width - kw + 1
# im2col: (batch, in_channels*kh*kw, out_h*out_w)
input_unfolded = F.unfold(input_tensor, kernel_size=(kh, kw), stride=1, padding=0)
# kernel: (out_channels, in_channels*kh*kw)
kernel_matrix = kernel.view(out_channels, -1)
# matmul broadcast: (out_channels, K) x (batch, K, out_h*out_w) -> (batch, out_channels, out_h*out_w)
result = torch.matmul(kernel_matrix, input_unfolded)
output.copy_(result.view(batch, out_channels, out_h, out_w))
return outputSource code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON