Skip to content
KernelIndex
Search⌘K

submission 551534

idanbeck · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 29 lines, June 9 Researcher Reciprocity License v1.0.

v_compile.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-grayscale-v2-551534?include=source"
interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
RGB to grayscalesuite of 6 cases
NVIDIA A100
8.10ms
#51 of 137
2026-03-14

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:16dad2233a0f9489f986309efc278f2711bff2b2fd45b1f9f69ee5561077e355
license declaredunknown
license concludedunknown
authorsidanbeck
imported2026-08-15

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

autotunereturn torch.compile(f, fullgraph=True, mode='max-autotune')

Kernel source

v_compile.py29 lines
#!POPCORN leaderboard grayscale_v2
#!POPCORN gpu A100
import torch
_W=None
_F=None

def _w(device,dtype):
    global _W
    if _W is None or _W.device!=device or _W.dtype!=dtype:
        _W=torch.tensor([0.2989,0.5870,0.1140],device=device,dtype=dtype)
    return _W

def _build(dtype):
    def f(x):
        return x[...,0]*0.2989 + x[...,1]*0.5870 + x[...,2]*0.1140
    try:
        return torch.compile(f, fullgraph=True, mode='max-autotune')
    except Exception:
        return f

@torch.inference_mode()
def custom_kernel(data):
    global _F
    rgb,out=data
    if _F is None:
        _F = _build(rgb.dtype)
    out[...] = _F(rgb)
    return out
scrolls · 29 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Changes from previous submission

Against this author's previous submission submission 551260.

#!POPCORN leaderboard grayscale_v2
#!POPCORN gpu A100
-
import torch
+ _W=None
+ _F=None
- _WEIGHTS = None
+ def _w(device,dtype):
+ global _W
+ if _W is None or _W.device!=device or _W.dtype!=dtype:
+ _W=torch.tensor([0.2989,0.5870,0.1140],device=device,dtype=dtype)
+ return _W
+ def _build(dtype):
+ def f(x):
+ return x[...,0]*0.2989 + x[...,1]*0.5870 + x[...,2]*0.1140
+ try:
+ return torch.compile(f, fullgraph=True, mode='max-autotune')
+ except Exception:
+ return f
- def _weights(device: torch.device, dtype: torch.dtype) -> torch.Tensor:
- global _WEIGHTS
- if _WEIGHTS is None or _WEIGHTS.device != device or _WEIGHTS.dtype != dtype:
- _WEIGHTS = torch.tensor([0.2989, 0.5870, 0.1140], device=device, dtype=dtype)
- return _WEIGHTS
-
-
@torch.inference_mode()
def custom_kernel(data):
- rgb, out = data
- # Reference-equivalent weighted sum over channel dimension.
- w = _weights(rgb.device, rgb.dtype)
- gray = torch.sum(rgb * w, dim=-1)
- out[...] = gray
+ global _F
+ rgb,out=data
+ if _F is None:
+ _F = _build(rgb.dtype)
+ out[...] = _F(rgb)
return out
scrolls · 42 diff lines total

Best evidence level for this revision: reported

JSON