submission 489475
jackkhuu · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 112 lines, June 9 Researcher Reciprocity License v1.0.
grayscale_py_H100_gpt-5-2_ka_submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-grayscale-v2-489475?include=source"interfacepython
Compatibility
measured onNVIDIA H100
declared hardwareNVIDIA H100
architecturessm_90
dtypesfp32
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:043afddbadf4e606275071fba54a0a497678e0cf30cf91be5e36c84af49f2ab4
license declaredunknown
license concludedunknown
authorsjackkhuu
imported2026-08-15
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
num-warps = 4
num_warps=4,Kernel source
grayscale_py_H100_gpt-5-2_ka_submission.py112 lines
# kernel.py
"""
RGB -> Grayscale (luma) conversion implemented as a single fused Triton kernel.
Fused stages (single pass, no intermediate tensors):
1) Load RGB (float32) from input image x[H, W, 3]
2) Compute Y = 0.2989 * R + 0.5870 * G + 0.1140 * B in fp32
3) Store grayscale output to y[H, W] (float32)
Wrapper restrictions respected:
- Wrapper only validates arguments, allocates output (optional), computes grid, and launches Triton.
- No PyTorch math ops are used for the computation.
"""
import torch
import triton
import triton.language as tl
@triton.jit
def _rgb_to_gray_kernel(
x_ptr, # *fp32, flattened view of (H*W*3)
y_ptr, # *fp32, flattened view of (H*W)
n_pixels, # int32/int64
BLOCK_SIZE: tl.constexpr,
):
pid = tl.program_id(axis=0)
pix = pid * BLOCK_SIZE + tl.arange(0, BLOCK_SIZE)
mask = pix < n_pixels
# x is contiguous (H, W, 3) => pixel p starts at p*3
base = pix * 3
r = tl.load(x_ptr + base + 0, mask=mask, other=0.0).to(tl.float32)
g = tl.load(x_ptr + base + 1, mask=mask, other=0.0).to(tl.float32)
b = tl.load(x_ptr + base + 2, mask=mask, other=0.0).to(tl.float32)
y = 0.2989 * r + 0.5870 * g + 0.1140 * b
tl.store(y_ptr + pix, y, mask=mask)
def kernel_function(x: torch.Tensor, y_out: torch.Tensor | None = None) -> torch.Tensor:
"""
Convert an RGB image to grayscale using a Triton kernel.
Args:
x: Tensor of shape (H, W, 3), dtype=float32, CUDA, contiguous.
y_out: Optional preallocated output tensor of shape (H, W), dtype=float32, CUDA, contiguous.
Returns:
y: Tensor of shape (H, W), dtype=float32 on CUDA.
"""
if not isinstance(x, torch.Tensor):
raise TypeError("x must be a torch.Tensor")
if not x.is_cuda:
raise ValueError("x must be on CUDA")
if x.dtype != torch.float32:
raise ValueError("x must be torch.float32")
if x.ndim != 3 or x.shape[-1] != 3:
raise ValueError("x must have shape (H, W, 3)")
if not x.is_contiguous():
raise ValueError("x must be contiguous")
H, W, C = x.shape
# Test guarantees square/even sizes, but kernel works for any H, W.
n_pixels = H * W
if y_out is None:
y = torch.empty((H, W), device=x.device, dtype=torch.float32)
else:
if not isinstance(y_out, torch.Tensor):
raise TypeError("y_out must be a torch.Tensor or None")
if not y_out.is_cuda or y_out.device != x.device:
raise ValueError("y_out must be on the same CUDA device as x")
if y_out.dtype != torch.float32:
raise ValueError("y_out must be torch.float32")
if tuple(y_out.shape) != (H, W):
raise ValueError(f"y_out must have shape {(H, W)}")
if not y_out.is_contiguous():
raise ValueError("y_out must be contiguous")
y = y_out
# Launch configuration
BLOCK_SIZE = 1024
grid = (triton.cdiv(n_pixels, BLOCK_SIZE),)
_rgb_to_gray_kernel[grid](
x, # pointers are passed directly; kernel treats them as flattened
y,
n_pixels,
BLOCK_SIZE=BLOCK_SIZE,
num_warps=4,
)
return y
import inspect
def custom_kernel(input):
sig = inspect.signature(kernel_function)
num_params = len(sig.parameters)
if len(input) == num_params:
return kernel_function(*input)
return kernel_function(input)
# Ensure deterministic cuBLAS.
import os
if os.environ.get("CUBLAS_WORKSPACE_CONFIG", "") not in (":4096:8", ":16:8"):
os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"
scrolls · 112 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON