submission 67964
rex_cz · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 83 lines, June 9 Researcher Reciprocity License v1.0.
grayscale_v4.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-grayscale-v2-67964?include=source"interfacepython
Compatibility
measured onNVIDIA H100
declared hardwareNVIDIA H100
architecturessm_90
dtypesfp32
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:ae76e2b9f839d3f99537625368fd99f73afb1cd1e01882eff85f10e3b6dc5233
license declaredunknown
license concludedunknown
authorsrex_cz
imported2026-08-15
Kernel source
grayscale_v4.py83 lines
# Reference: https://github.com/gpu-mode/reference-kernels/blob/main/problems/pmpp_v2/grayscale_py/reference.py
# B200: block dim 256: 64381.682μs
# B200: block dim 1024: 48362.844μs
import torch
from task import input_t, output_t
import cutlass.cute as cute
from cutlass.cute.runtime import from_dlpack
@cute.kernel
def grayscale_kernel(
gA: cute.Tensor,
gC: cute.Tensor,
):
tidx, _, _ = cute.arch.thread_idx()
bidx, _, _ = cute.arch.block_idx()
bdimx, _, _ = cute.arch.block_dim()
m, n, _ = gA.shape[1]
thread_idx = bidx * bdimx + tidx
mi = thread_idx // n
ni = thread_idx % n
# rgb.shape: ((1,4),3)
rgb = gA[(None, (mi, ni, None))].load()
r, g, b = rgb[(None, 0)], rgb[(None, 1)], rgb[(None, 2)]
y = 0.2989 * r + 0.5870 * g + 0.1140 * b
gC[(None, (mi, ni))].store(y)
@cute.jit
def grayscale(input):
mA, mC = input
num_threads_per_block = 1024
# gA.shape: ((1,4),(128,32,3))
gA = cute.zipped_divide(mA, (1, 4))
gC = cute.zipped_divide(mC, (1, 4))
grayscale_kernel(gA, gC).launch(
grid=[cute.size(gA, mode=[1]) // num_threads_per_block // 3, 1, 1],
block=[num_threads_per_block, 1, 1],
)
def generate_input(size: int, seed: int) -> input_t:
"""
Generates random RGB image tensor of specified size.
Returns:
Tensor of shape (size, size, 3) with values in [0, 1]
"""
gen = torch.Generator(device="cuda")
gen.manual_seed(seed)
x = torch.rand(
size, size, 3, device="cuda", dtype=torch.float32, generator=gen
).contiguous()
y = torch.empty(size, size, device="cuda", dtype=torch.float32).contiguous()
return x, y
sizes = [
(1001, 128),
(5531, 256),
(9173, 512),
(93246, 1024),
(6256, 2048),
(8841, 4096),
(6252, 8192),
(54352, 16384),
]
compiled_kernels = {}
for i in sizes:
seed, size = i
data, output = generate_input(size, seed)
a_ = from_dlpack(data, assumed_align=16)
c_ = from_dlpack(output, assumed_align=16)
compiled_kernels[size] = cute.compile(grayscale, (a_, c_))
def custom_kernel(input):
compiled_kernels[input[0].shape[0]](input)
return input[1]
scrolls · 83 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Changes from previous submission
Against this author's previous submission submission 67954.
Best evidence level for this revision: reported
JSON