Skip to content
KernelIndex
Search⌘K

submission 67528

yue · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 72 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-grayscale-v2-67528?include=source"
interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
RGB to grayscalesuite of 6 cases
NVIDIA A100
8.66ms
#52 of 137
2025-11-07

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:798ff5eeba5e84fcc9d1159429c18fc41ca5da9e3c69dbb9cf72ff8de2403101
license declaredunknown
license concludedunknown
authorsyue
imported2026-08-15

Kernel source

submission.py72 lines
#!POPCORN leaderboard grayscale_v2

from task import input_t, output_t
from utils import DeterministicContext
import triton
import triton.language as tl


@triton.jit
def _grayscale_kernel(
    rgb_ptr,
    output_ptr,
    height: tl.int32,
    width: tl.int32,
    stride_h: tl.int32,
    stride_w: tl.int32,
    stride_c: tl.int32,
    out_stride_h: tl.int32,
    out_stride_w: tl.int32,
    BLOCK_H: tl.constexpr,
    BLOCK_W: tl.constexpr,
):
    pid_h = tl.program_id(0)
    pid_w = tl.program_id(1)

    offs_h = pid_h * BLOCK_H + tl.arange(0, BLOCK_H)
    offs_w = pid_w * BLOCK_W + tl.arange(0, BLOCK_W)

    mask_h = offs_h < height
    mask_w = offs_w < width

    hh = offs_h[:, None]
    ww = offs_w[None, :]
    mask = mask_h[:, None] & mask_w[None, :]

    base = hh * stride_h + ww * stride_w

    r = tl.load(rgb_ptr + base + 0 * stride_c, mask=mask, other=0.0)
    g = tl.load(rgb_ptr + base + 1 * stride_c, mask=mask, other=0.0)
    b = tl.load(rgb_ptr + base + 2 * stride_c, mask=mask, other=0.0)

    grayscale = 0.2989 * r + 0.5870 * g + 0.1140 * b

    tl.store(output_ptr + hh * out_stride_h + ww * out_stride_w, grayscale, mask=mask)


def custom_kernel(data: input_t) -> output_t:
    with DeterministicContext():
        rgb, output = data
        height, width, channels = rgb.shape
        if channels != 3:
            raise ValueError(f"Expected last dimension to be 3, got {channels}")

        BLOCK_H = 32
        BLOCK_W = 32

        grid = (triton.cdiv(height, BLOCK_H), triton.cdiv(width, BLOCK_W))

        _grayscale_kernel[grid](
            rgb,
            output,
            height,
            width,
            rgb.stride(0),
            rgb.stride(1),
            rgb.stride(2),
            output.stride(0),
            output.stride(1),
            BLOCK_H=BLOCK_H,
            BLOCK_W=BLOCK_W,
        )
        return output
scrolls · 72 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON