submission 551534
idanbeck · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 29 lines, June 9 Researcher Reciprocity License v1.0.
v_compile.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-grayscale-v2-551534?include=source"interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp32
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:16dad2233a0f9489f986309efc278f2711bff2b2fd45b1f9f69ee5561077e355
license declaredunknown
license concludedunknown
authorsidanbeck
imported2026-08-15
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
autotune
return torch.compile(f, fullgraph=True, mode='max-autotune')Kernel source
v_compile.py29 lines
#!POPCORN leaderboard grayscale_v2
#!POPCORN gpu A100
import torch
_W=None
_F=None
def _w(device,dtype):
global _W
if _W is None or _W.device!=device or _W.dtype!=dtype:
_W=torch.tensor([0.2989,0.5870,0.1140],device=device,dtype=dtype)
return _W
def _build(dtype):
def f(x):
return x[...,0]*0.2989 + x[...,1]*0.5870 + x[...,2]*0.1140
try:
return torch.compile(f, fullgraph=True, mode='max-autotune')
except Exception:
return f
@torch.inference_mode()
def custom_kernel(data):
global _F
rgb,out=data
if _F is None:
_F = _build(rgb.dtype)
out[...] = _F(rgb)
return out
scrolls · 29 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Changes from previous submission
Against this author's previous submission submission 551260.
#!POPCORN leaderboard grayscale_v2#!POPCORN gpu A100-import torch+ _W=None+ _F=None- _WEIGHTS = None+ def _w(device,dtype):+ global _W+ if _W is None or _W.device!=device or _W.dtype!=dtype:+ _W=torch.tensor([0.2989,0.5870,0.1140],device=device,dtype=dtype)+ return _W+ def _build(dtype):+ def f(x):+ return x[...,0]*0.2989 + x[...,1]*0.5870 + x[...,2]*0.1140+ try:+ return torch.compile(f, fullgraph=True, mode='max-autotune')+ except Exception:+ return f- def _weights(device: torch.device, dtype: torch.dtype) -> torch.Tensor:- global _WEIGHTS- if _WEIGHTS is None or _WEIGHTS.device != device or _WEIGHTS.dtype != dtype:- _WEIGHTS = torch.tensor([0.2989, 0.5870, 0.1140], device=device, dtype=dtype)- return _WEIGHTS--@torch.inference_mode()def custom_kernel(data):- rgb, out = data- # Reference-equivalent weighted sum over channel dimension.- w = _weights(rgb.device, rgb.dtype)- gray = torch.sum(rgb * w, dim=-1)- out[...] = gray+ global _F+ rgb,out=data+ if _F is None:+ _F = _build(rgb.dtype)+ out[...] = _F(rgb)return out
scrolls · 42 diff lines total
Best evidence level for this revision: reported
JSON