Skip to content
KernelIndex
Search⌘K

submission 782511

ooousay · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 17 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-conv2d-v2-782511?include=source"
interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
2D convolutionsuite of 5 cases
NVIDIA A100
2.66s
#40 of 40
2026-05-12

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:8628e34796163482e0be8cdab6473d5e0b4f05686b98a14620bbc96acde0d5cd
license declaredunknown
license concludedunknown
authorsooousay
imported2026-08-15

Kernel source

submission.py17 lines
"""Current best submission. Starts as the cuDNN baseline.

TF32 must stay off: the reference runs in DeterministicContext (strict FP32),
and default-TF32 cuDNN diverges by >1e-3 on the K=32 shape.
"""
from task import input_t, output_t
import torch
import torch.nn.functional as F

torch.backends.cudnn.allow_tf32 = False


def custom_kernel(data: input_t) -> output_t:
    input_tensor, kernel, output = data
    output[...] = F.conv2d(input_tensor, kernel, stride=1, padding=0)
    return output

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON