Skip to content
KernelIndex
Search⌘K

submission 551659

idanbeck · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 22 lines, June 9 Researcher Reciprocity License v1.0.

compiled_out_fn.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-grayscale-v2-551659?include=source"
interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
RGB to grayscalesuite of 6 cases
NVIDIA A100
2.55ms
#27 of 137
2026-03-14

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:ba7c547dec5e89737c7f97f25dfd257b36936f2c1adde0c21331f23ab87dfda3
license declaredunknown
license concludedunknown
authorsidanbeck
imported2026-08-15

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

autotunereturn torch.compile(f, fullgraph=False, mode='max-autotune-no-cudagraphs')

Kernel source

compiled_out_fn.py22 lines
#!POPCORN leaderboard grayscale_v2
#!POPCORN gpu A100
import torch
_F=None

def _build():
    def f(x, out):
        out[...] = x[...,0]*0.2989 + x[...,1]*0.5870 + x[...,2]*0.1140
        return out
    try:
        return torch.compile(f, fullgraph=False, mode='max-autotune-no-cudagraphs')
    except Exception:
        return f

@torch.inference_mode()
def custom_kernel(data):
    global _F
    rgb,out=data
    if _F is None:
        _F=_build()
    return _F(rgb, out)

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Changes from previous submission

Against this author's previous submission submission 551588.

#!POPCORN leaderboard grayscale_v2
#!POPCORN gpu A100
import torch
- _F = None
+ _F=None
-
def _build():
- def f(x):
- return x[..., 0] * 0.2989 + x[..., 1] * 0.5870 + x[..., 2] * 0.1140
-
+ def f(x, out):
+ out[...] = x[...,0]*0.2989 + x[...,1]*0.5870 + x[...,2]*0.1140
+ return out
try:
- return torch.compile(f, fullgraph=False, mode="max-autotune-no-cudagraphs")
+ return torch.compile(f, fullgraph=False, mode='max-autotune-no-cudagraphs')
except Exception:
return f
-
@torch.inference_mode()
def custom_kernel(data):
global _F
- rgb, out = data
+ rgb,out=data
if _F is None:
- _F = _build()
- out[...] = _F(rgb)
- return out
+ _F=_build()
+ return _F(rgb, out)
scrolls · 32 diff lines total

Best evidence level for this revision: reported

JSON