Skip to content
KernelIndex
Search⌘K

submission 782512

ooousay · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 36 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-conv2d-v2-782512?include=source"
interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
2D convolutionsuite of 5 cases
NVIDIA A100
19.5ms
#13 of 40
2026-05-13

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:02cb3f02f7d4351e6926b06d2bb11fa54c2334e017be87dde0fe89869e589654
license declaredunknown
license concludedunknown
authorsooousay
imported2026-08-15

Kernel source

submission.py36 lines
"""Current best submission.

cuDNN baseline + `cudnn.benchmark = True`.

Root cause of the original 2.59 s outlier on S256 K32 C128 B1: with
`cudnn.deterministic = True` (set by the reference's DeterministicContext) and
`benchmark = False` (PyTorch default), cuDNN picks an FFT-conv engine that
degenerates to a memory-bound mat-vec at B=1.

Flipping `benchmark = True` lets cuDNN run a timed heuristic on the first call
and select a much faster implicit-GEMM path for K=32. The eval harness does
one warmup call before timing, so algo selection happens off-clock.

Tradeoff: `benchmark = True` can pick worse algos for some mid-sized shapes
when L2 is cleared between calls (the heuristic times candidates with hot L2).
On this benchmark suite, that costs ~30 ms across all shapes vs cuDNN-default
but saves ~2470 ms on the K=32 shape — net ~11x speedup on total wall time.

TF32 stays off: reference is strict FP32, default-TF32 cuDNN diverges >1e-3
on K=16 and K=32 shapes.
"""
from task import input_t, output_t
import torch
import torch.nn.functional as F

torch.backends.cudnn.allow_tf32 = False
torch.backends.cuda.matmul.allow_tf32 = False
torch.backends.cudnn.benchmark = True
torch.backends.cudnn.deterministic = False


def custom_kernel(data: input_t) -> output_t:
    input_tensor, kernel, output = data
    output[...] = F.conv2d(input_tensor, kernel, stride=1, padding=0)
    return output
scrolls · 36 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Changes from previous submission

Against this author's previous submission submission 782511.

- """Current best submission. Starts as the cuDNN baseline.
+ """Current best submission.
- TF32 must stay off: the reference runs in DeterministicContext (strict FP32),
- and default-TF32 cuDNN diverges by >1e-3 on the K=32 shape.
+ cuDNN baseline + `cudnn.benchmark = True`.
+
+ Root cause of the original 2.59 s outlier on S256 K32 C128 B1: with
+ `cudnn.deterministic = True` (set by the reference's DeterministicContext) and
+ `benchmark = False` (PyTorch default), cuDNN picks an FFT-conv engine that
+ degenerates to a memory-bound mat-vec at B=1.
+
+ Flipping `benchmark = True` lets cuDNN run a timed heuristic on the first call
+ and select a much faster implicit-GEMM path for K=32. The eval harness does
+ one warmup call before timing, so algo selection happens off-clock.
+
+ Tradeoff: `benchmark = True` can pick worse algos for some mid-sized shapes
+ when L2 is cleared between calls (the heuristic times candidates with hot L2).
+ On this benchmark suite, that costs ~30 ms across all shapes vs cuDNN-default
+ but saves ~2470 ms on the K=32 shape — net ~11x speedup on total wall time.
+
+ TF32 stays off: reference is strict FP32, default-TF32 cuDNN diverges >1e-3
+ on K=16 and K=32 shapes.
"""
from task import input_t, output_t
import torch
import torch.nn.functional as F
torch.backends.cudnn.allow_tf32 = False
+ torch.backends.cuda.matmul.allow_tf32 = False
+ torch.backends.cudnn.benchmark = True
+ torch.backends.cudnn.deterministic = False
def custom_kernel(data: input_t) -> output_t:
scrolls · 35 diff lines total

Best evidence level for this revision: reported

JSON