Skip to content
KernelIndex
Search⌘K

submission 780434

Kernel-Zhang · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 85 lines, June 9 Researcher Reciprocity License v1.0.

ref.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-conv2d-v2-780434?include=source"
interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
2D convolutionsuite of 5 cases
NVIDIA A100
2.66s
#39 of 40
2026-04-28

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:62ca5c72e3034d1dfcfcb3d55f40b07bc1af4e878225cf3e450dbfb7d5a4e042
license declaredunknown
license concludedunknown
authorsKernel-Zhang
imported2026-08-15

Kernel source

ref.py85 lines
from utils import make_match_reference, DeterministicContext
import torch
import torch.nn.functional as F
from task import input_t, output_t

def custom_kernel(data: input_t) -> output_t:
    """
    Reference implementation of 2D convolution using PyTorch.
    Args:
        data: Tuple of (input tensor, kernel tensor)
    Returns:
        Output tensor after convolution
    """
    with DeterministicContext():
        input_tensor, kernel, output = data
        return F.conv2d(
            input_tensor,
            kernel,
            # No padding and no striding
            stride=1,
            padding=0,
        )

def ref_kernel(data: input_t) -> output_t:
    """
    Reference implementation of 2D convolution using PyTorch.
    Args:
        data: Tuple of (input tensor, kernel tensor)
    Returns:
        Output tensor after convolution
    """
    with DeterministicContext():
        input_tensor, kernel, output = data
        return F.conv2d(
            input_tensor,
            kernel,
            # No padding and no striding
            stride=1,
            padding=0,
        )


def generate_input(
    size: int, kernelsize: int, channels: int, batch: int, seed: int
) -> input_t:
    """
    Generates random input and kernel tensors.
    Returns:
        Tuple of (input tensor, kernel tensor)
    """
    gen = torch.Generator(device="cuda")
    gen.manual_seed(seed)

    # Generate input tensor: [batch, in_channels, height, width]
    input_tensor = torch.randn(
        batch, channels, size, size, device="cuda", dtype=torch.float32, generator=gen
    ).contiguous()

    # Generate kernel tensor: [out_channels, in_channels, kernel_height, kernel_width]
    # Here we use same number of output channels as input channels for simplicity
    kernel = torch.randn(
        channels,
        channels,
        kernelsize,
        kernelsize,
        device="cuda",
        dtype=torch.float32,
        generator=gen,
    ).contiguous()

    output_tensor = torch.empty(
        batch,
        channels,
        size - kernelsize + 1,
        size - kernelsize + 1,
        device="cuda",
        dtype=torch.float32,
    )

    return input_tensor, kernel, output_tensor


check_implementation = make_match_reference(ref_kernel, rtol=1e-3, atol=1e-3)

scrolls · 85 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON