Skip to content
KernelIndex
Search⌘K

submission 43581

davidberard · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 24 lines, June 9 Researcher Reciprocity License v1.0.

triton.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-grayscale-v2-43581?include=source"
interfacepython
Compatibility
measured onNVIDIA H100
declared hardwareNVIDIA H100
architecturessm_90
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
RGB to grayscalesuite of 6 cases
NVIDIA H100
1.42ms
#31 of 36
2025-09-24

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:2f37eb3b764240f0dd58a60a18e7da1ff8f4d5789174a533cd55dbb1ff50eeca
license declaredunknown
license concludedunknown
authorsdavidberard
imported2026-08-15

Kernel source

triton.py24 lines
import torch
import triton
import triton.language as tl
#!POPCORN leaderboard identity_py
from task import input_t, output_t

@triton.jit
def kernel(inp, out, numel, BLOCK_SIZE: tl.constexpr):
    pid = tl.program_id(0)
    offs = tl.arange(0, BLOCK_SIZE) + pid * BLOCK_SIZE
    red = tl.load(inp + offs * 3, mask=(offs * 3< numel))
    green = tl.load(inp + offs * 3 + 1, mask=(offs * 3 + 1 < numel))
    blue = tl.load(inp + offs * 3 + 2, mask=(offs * 3 + 2 < numel))
    res = 0.2989 * red + 0.5870 * green + 0.1140 * blue
    tl.store(out + offs, res, mask=(offs * 3 < numel))

# User kernel implementation.
def custom_kernel(input: input_t) -> output_t:
    data, output = input
    numel = data.numel()
    BLOCK_SIZE = 128
    grid = (triton.cdiv(numel, 3 * BLOCK_SIZE),)
    kernel[grid](data, output, numel, BLOCK_SIZE)
    return output

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON