Skip to content
KernelIndex
Search⌘K

submission 43583

davidberard · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 25 lines, June 9 Researcher Reciprocity License v1.0.

triton.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-grayscale-v2-43583?include=source"
interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
RGB to grayscalesuite of 6 cases
NVIDIA A100
2.50ms
#20 of 137
2025-09-25

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:e107fa90deb1d81412f2c525acd55af1064c6e188e0ee06ebf8193760edabbe1
license declaredunknown
license concludedunknown
authorsdavidberard
imported2026-08-15

Kernel source

triton.py25 lines
import torch
import triton
import triton.language as tl
#!POPCORN leaderboard identity_py
from task import input_t, output_t

@triton.jit
def kernel(inp, out, numel, BLOCK_SIZE: tl.constexpr):
    pid = tl.program_id(0)
    offs = tl.arange(0, BLOCK_SIZE) + pid * BLOCK_SIZE
    tl.arange(0, 4)
    red = tl.load(inp + offs * 3, mask=(offs * 3 < numel))
    green = tl.load(inp + offs * 3 + 1, mask=(offs * 3 + 1 < numel))
    blue = tl.load(inp + offs * 3 + 2, mask=(offs * 3 + 2 < numel))
    res = 0.2989 * red + 0.5870 * green + 0.1140 * blue
    tl.store(out + offs, res, mask=(offs * 3 < numel))

# User kernel implementation.
def custom_kernel(input: input_t) -> output_t:
    data, output = input
    numel = data.numel()
    BLOCK_SIZE = 128
    grid = (triton.cdiv(numel, 3 * BLOCK_SIZE),)
    kernel[grid](data, output, numel, BLOCK_SIZE)
    return output

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Changes from previous submission

Against this author's previous submission submission 43581.

⋯ 7 unchanged lines
def kernel(inp, out, numel, BLOCK_SIZE: tl.constexpr):
pid = tl.program_id(0)
offs = tl.arange(0, BLOCK_SIZE) + pid * BLOCK_SIZE
- red = tl.load(inp + offs * 3, mask=(offs * 3< numel))
+ tl.arange(0, 4)
+ red = tl.load(inp + offs * 3, mask=(offs * 3 < numel))
green = tl.load(inp + offs * 3 + 1, mask=(offs * 3 + 1 < numel))
blue = tl.load(inp + offs * 3 + 2, mask=(offs * 3 + 2 < numel))
res = 0.2989 * red + 0.5870 * green + 0.1140 * blue

Best evidence level for this revision: reported

JSON