submission 551567
idanbeck · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 25 lines, June 9 Researcher Reciprocity License v1.0.
submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-grayscale-v2-551567?include=source"interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp32
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:ee93c7228ab95ae0c1160c5dce3631aeb69602e0207c25e9fb243aff769368a2
license declaredunknown
license concludedunknown
authorsidanbeck
imported2026-08-15
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
autotune
return torch.compile(f, fullgraph=True, mode="max-autotune")Kernel source
submission.py25 lines
#!POPCORN leaderboard grayscale_v2
#!POPCORN gpu A100
import torch
_F = None
def _build():
def f(x):
return x[..., 0] * 0.2989 + x[..., 1] * 0.5870 + x[..., 2] * 0.1140
try:
return torch.compile(f, fullgraph=True, mode="max-autotune")
except Exception:
return f
@torch.inference_mode()
def custom_kernel(data):
global _F
rgb, out = data
if _F is None:
_F = _build()
out[...] = _F(rgb)
return out
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Changes from previous submission
Against this author's previous submission submission 551534.
#!POPCORN leaderboard grayscale_v2#!POPCORN gpu A100import torch- _W=None- _F=None+ _F = None- def _w(device,dtype):- global _W- if _W is None or _W.device!=device or _W.dtype!=dtype:- _W=torch.tensor([0.2989,0.5870,0.1140],device=device,dtype=dtype)- return _W- def _build(dtype):+ def _build():def f(x):- return x[...,0]*0.2989 + x[...,1]*0.5870 + x[...,2]*0.1140+ return x[..., 0] * 0.2989 + x[..., 1] * 0.5870 + x[..., 2] * 0.1140+try:- return torch.compile(f, fullgraph=True, mode='max-autotune')+ return torch.compile(f, fullgraph=True, mode="max-autotune")except Exception:return f+@torch.inference_mode()def custom_kernel(data):global _F- rgb,out=data+ rgb, out = dataif _F is None:- _F = _build(rgb.dtype)+ _F = _build()out[...] = _F(rgb)return out
scrolls · 36 diff lines total
Best evidence level for this revision: reported
JSON