Skip to content
KernelIndex
Search⌘K

submission 525090

krasnaya_66854 · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 55 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-conv2d-v2-525090?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
2D convolutionsuite of 5 cases
NVIDIA B200
51.9ms
#18 of 28
2026-03-10

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:42208061ebae51fc02be9d53902bae9e2b9561dc880aa888cccc4a0dbe7ad461
license declaredunknown
license concludedunknown
authorskrasnaya_66854
imported2026-08-15

Kernel source

submission.py55 lines
import os
os.environ['CUBLAS_WORKSPACE_CONFIG'] = ':4096:8'

import torch
import torch.nn.functional as F
from task import input_t, output_t
from utils import DeterministicContext


def _next_fast_len(n: int) -> int:
    """Smallest integer >= n with only prime factors 2, 3, 5 (5-smooth = fast for cuFFT)."""
    while True:
        m = n
        for p in (2, 3, 5):
            while m % p == 0:
                m //= p
        if m == 1:
            return n
        n += 1


def _fft_conv2d(input_tensor: torch.Tensor, kernel: torch.Tensor, output: torch.Tensor) -> torch.Tensor:
    B, Ci, H, W = input_tensor.shape
    Co, _, kH, kW = kernel.shape
    H_out, W_out = H - kH + 1, W - kW + 1

    fH = _next_fast_len(H + kH - 1)
    fW = _next_fast_len(W + kW - 1)
    fW2 = fW // 2 + 1

    X_fft = torch.fft.rfft2(input_tensor, s=(fH, fW))
    W_fft = torch.fft.rfft2(
        kernel.reshape(Co * Ci, kH, kW), s=(fH, fW)
    ).reshape(Co, Ci, fH, fW2)

    out_fft = torch.einsum('bihw,oihw->bohw', X_fft, W_fft.conj())
    output[...] = torch.fft.irfft2(out_fft, s=(fH, fW))[:, :, :H_out, :W_out]
    return output


def custom_kernel(data: input_t) -> output_t:
    input_tensor, kernel, output = data
    Co, Ci, kH, kW = kernel.shape

    if kW <= 16:
        # FFT convolution. For k<=16, FFT error is ~2e-5 (well within 1e-3 tolerance).
        # W_fft is at most (128^2) * 288 * 145 * 8 = 5.1 GB which fits in A100 80GB.
        return _fft_conv2d(input_tensor, kernel, output)
    else:
        # k=32: FFT float32 error (~4e-3) exceeds atol=1e-3 for near-zero elements.
        # Must use DeterministicContext to reproduce the reference's exact float32 result.
        with DeterministicContext():
            output[...] = F.conv2d(input_tensor, kernel, stride=1, padding=0)
        return output
scrolls · 55 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Changes from previous submission

Against this author's previous submission submission 525086.

Best evidence level for this revision: reported

JSON