Skip to content
KernelIndex
Search⌘K

submission 489475

jackkhuu · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 112 lines, June 9 Researcher Reciprocity License v1.0.

grayscale_py_H100_gpt-5-2_ka_submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-grayscale-v2-489475?include=source"
interfacepython
Compatibility
measured onNVIDIA H100
declared hardwareNVIDIA H100
architecturessm_90
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
RGB to grayscalesuite of 6 cases
NVIDIA H100
1.39ms
#30 of 36
2026-02-12

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:043afddbadf4e606275071fba54a0a497678e0cf30cf91be5e36c84af49f2ab4
license declaredunknown
license concludedunknown
authorsjackkhuu
imported2026-08-15

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

num-warps = 4num_warps=4,

Kernel source

grayscale_py_H100_gpt-5-2_ka_submission.py112 lines
# kernel.py
"""
RGB -> Grayscale (luma) conversion implemented as a single fused Triton kernel.

Fused stages (single pass, no intermediate tensors):
  1) Load RGB (float32) from input image x[H, W, 3]
  2) Compute Y = 0.2989 * R + 0.5870 * G + 0.1140 * B in fp32
  3) Store grayscale output to y[H, W] (float32)

Wrapper restrictions respected:
  - Wrapper only validates arguments, allocates output (optional), computes grid, and launches Triton.
  - No PyTorch math ops are used for the computation.
"""

import torch
import triton
import triton.language as tl


@triton.jit
def _rgb_to_gray_kernel(
    x_ptr,  # *fp32, flattened view of (H*W*3)
    y_ptr,  # *fp32, flattened view of (H*W)
    n_pixels,  # int32/int64
    BLOCK_SIZE: tl.constexpr,
):
    pid = tl.program_id(axis=0)
    pix = pid * BLOCK_SIZE + tl.arange(0, BLOCK_SIZE)
    mask = pix < n_pixels

    # x is contiguous (H, W, 3) => pixel p starts at p*3
    base = pix * 3

    r = tl.load(x_ptr + base + 0, mask=mask, other=0.0).to(tl.float32)
    g = tl.load(x_ptr + base + 1, mask=mask, other=0.0).to(tl.float32)
    b = tl.load(x_ptr + base + 2, mask=mask, other=0.0).to(tl.float32)

    y = 0.2989 * r + 0.5870 * g + 0.1140 * b
    tl.store(y_ptr + pix, y, mask=mask)


def kernel_function(x: torch.Tensor, y_out: torch.Tensor | None = None) -> torch.Tensor:
    """
    Convert an RGB image to grayscale using a Triton kernel.

    Args:
        x: Tensor of shape (H, W, 3), dtype=float32, CUDA, contiguous.
        y_out: Optional preallocated output tensor of shape (H, W), dtype=float32, CUDA, contiguous.

    Returns:
        y: Tensor of shape (H, W), dtype=float32 on CUDA.
    """
    if not isinstance(x, torch.Tensor):
        raise TypeError("x must be a torch.Tensor")
    if not x.is_cuda:
        raise ValueError("x must be on CUDA")
    if x.dtype != torch.float32:
        raise ValueError("x must be torch.float32")
    if x.ndim != 3 or x.shape[-1] != 3:
        raise ValueError("x must have shape (H, W, 3)")
    if not x.is_contiguous():
        raise ValueError("x must be contiguous")

    H, W, C = x.shape
    # Test guarantees square/even sizes, but kernel works for any H, W.
    n_pixels = H * W

    if y_out is None:
        y = torch.empty((H, W), device=x.device, dtype=torch.float32)
    else:
        if not isinstance(y_out, torch.Tensor):
            raise TypeError("y_out must be a torch.Tensor or None")
        if not y_out.is_cuda or y_out.device != x.device:
            raise ValueError("y_out must be on the same CUDA device as x")
        if y_out.dtype != torch.float32:
            raise ValueError("y_out must be torch.float32")
        if tuple(y_out.shape) != (H, W):
            raise ValueError(f"y_out must have shape {(H, W)}")
        if not y_out.is_contiguous():
            raise ValueError("y_out must be contiguous")
        y = y_out

    # Launch configuration
    BLOCK_SIZE = 1024
    grid = (triton.cdiv(n_pixels, BLOCK_SIZE),)

    _rgb_to_gray_kernel[grid](
        x,  # pointers are passed directly; kernel treats them as flattened
        y,
        n_pixels,
        BLOCK_SIZE=BLOCK_SIZE,
        num_warps=4,
    )
    return y

import inspect

def custom_kernel(input):
    sig = inspect.signature(kernel_function)
    num_params = len(sig.parameters)

    if len(input) == num_params:
        return kernel_function(*input)
    return kernel_function(input)


# Ensure deterministic cuBLAS.
import os
if os.environ.get("CUBLAS_WORKSPACE_CONFIG", "") not in (":4096:8", ":16:8"):
    os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"

scrolls · 112 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON