Skip to content
KernelIndex
Search⌘K

submission 512856

JordanNanos · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 25 lines, June 9 Researcher Reciprocity License v1.0.

conv2d_py_submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-conv2d-v2-512856?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
2D convolutionsuite of 5 cases
NVIDIA B200
42.0ms
#12 of 28
2026-03-05

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:4b337fcdf73f8b9487b2e7dd0c2e083bb7878c2c41a366cbe1a8a692cdc7cae9
license declaredunknown
license concludedunknown
authorsJordanNanos
imported2026-08-15

Kernel source

conv2d_py_submission.py25 lines
import torch
import torch.nn.functional as F
from task import input_t, output_t


def custom_kernel(data: input_t) -> output_t:
    input_tensor, kernel, output = data
    
    batch, channels, height, width = input_tensor.shape
    out_channels, in_channels, kh, kw = kernel.shape
    
    out_h = height - kh + 1
    out_w = width - kw + 1
    
    # im2col: (batch, in_channels*kh*kw, out_h*out_w)
    input_unfolded = F.unfold(input_tensor, kernel_size=(kh, kw), stride=1, padding=0)
    
    # kernel: (out_channels, in_channels*kh*kw)
    kernel_matrix = kernel.view(out_channels, -1)
    
    # matmul broadcast: (out_channels, K) x (batch, K, out_h*out_w) -> (batch, out_channels, out_h*out_w)
    result = torch.matmul(kernel_matrix, input_unfolded)
    
    output.copy_(result.view(batch, out_channels, out_h, out_w))
    return output

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON