submission 553311
idanbeck · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 27 lines, June 9 Researcher Reciprocity License v1.0.
best_bump.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-grayscale-v2-553311?include=source"interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp32
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:7c1f15b7f716c9fa1c6212b70ce161db0af3785d058084a6c629239d189c16b8
license declaredunknown
license concludedunknown
authorsidanbeck
imported2026-08-15
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
autotune
return torch.compile(f, fullgraph=False, mode='max-autotune-no-cudagraphs')num-warps = 4
num_warps = 4Kernel source
best_bump.py27 lines
#!POPCORN leaderboard grayscale_v2
#!POPCORN gpu A100
import torch
_F=None
def _build():
def f(x, out):
out[...] = x[...,0]*0.2989 + x[...,1]*0.5870 + x[...,2]*0.1140
return out
try:
return torch.compile(f, fullgraph=False, mode='max-autotune-no-cudagraphs')
except Exception:
return f
@torch.inference_mode()
def custom_kernel(data):
global _F
rgb,out=data
if _F is None:
_F=_build()
return _F(rgb, out)
# tge-template: generic cuda knobs
block_size = 256
num_warps = 4
unroll_factor = 2
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Changes from previous submission
Against this author's previous submission submission 551659.
⋯ 18 unchanged linesif _F is None:_F=_build()return _F(rgb, out)++ # tge-template: generic cuda knobs+ block_size = 256+ num_warps = 4+ unroll_factor = 2
Best evidence level for this revision: reported
JSON