Skip to content
KernelIndex
Search⌘K

submission 510194

ağaç.mp4 · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 83 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-grayscale-v2-510194?include=source"
interfacepython
Compatibility
measured onNVIDIA L4
declared hardwareNVIDIA L4
architecturessm_89
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
RGB to grayscalesuite of 6 cases
NVIDIA L4
17.2ms
#4 of 12
2026-02-26

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:90c0d320c9e40955d2e447c2d57eb01d33db7fff3692bdd1091d3fa502dbd38c
license declaredunknown
license concludedunknown
authorsağaç.mp4
imported2026-08-15

Techniques

Extracted from the mirrored source by pattern, never inferred. Each row cites its line.

vector-width = float4v6: 4px/thread, 256 threads, float4 loads+stores, explicit FMA

Kernel source

submission.py83 lines
"""
Grayscale — RGB to Grayscale conversion
@MemoryCoalesced

Y = 0.2989 R + 0.5870 G + 0.1140 B

Input:  (H, W, 3) float32
Output: (H, W)    float32

v6: 4px/thread, 256 threads, float4 loads+stores, explicit FMA
    22 registers, 0 spills — 89.8% bandwidth efficiency on B200
"""

import torch
from torch.utils.cpp_extension import load_inline

cpp_source = """
void grayscale_cuda(torch::Tensor input, torch::Tensor output);
"""

cuda_source = r"""
#include <torch/extension.h>
#include <cuda_runtime.h>

#define THREADS 128
#define CR 0.2989f
#define CG 0.5870f
#define CB 0.1140f

__global__ void __launch_bounds__(THREADS)
grayscale_kernel(
    const float4* __restrict__ input,
    float4*       __restrict__ output,
    int num_groups)
{
    const int idx = blockIdx.x * THREADS + threadIdx.x;
    if (idx >= num_groups) return;

    const float4 a = __ldg(&input[idx * 3 + 0]); // R0 G0 B0 R1
    const float4 b = __ldg(&input[idx * 3 + 1]); // G1 B1 R2 G2
    const float4 c = __ldg(&input[idx * 3 + 2]); // B2 R3 G3 B3

    float4 result;
    result.x = __fmaf_rn(CR, a.x, __fmaf_rn(CG, a.y, CB * a.z));
    result.y = __fmaf_rn(CR, a.w, __fmaf_rn(CG, b.x, CB * b.y));
    result.z = __fmaf_rn(CR, b.z, __fmaf_rn(CG, b.w, CB * c.x));
    result.w = __fmaf_rn(CR, c.y, __fmaf_rn(CG, c.z, CB * c.w));

    output[idx] = result;
}

void grayscale_cuda(torch::Tensor input, torch::Tensor output) {
    const int num_pixels = input.size(0) * input.size(1);
    const int num_groups = num_pixels / 4;
    const int blocks     = (num_groups + THREADS - 1) / THREADS;

    grayscale_kernel<<<blocks, THREADS>>>(
        reinterpret_cast<const float4*>(input.data_ptr<float>()),
        reinterpret_cast<float4*>(output.data_ptr<float>()),
        num_groups);
}
"""

module = None

def get_module():
    global module
    if module is None:
        module = load_inline(
            name='grayscale_v9',
            cpp_sources=cpp_source,
            cuda_sources=cuda_source,
            functions=['grayscale_cuda'],
            verbose=True,
            extra_cuda_cflags=['-O3', '--use_fast_math', '-std=c++17'],
        )
    return module

def custom_kernel(data: tuple) -> torch.Tensor:
    inp, out = data
    get_module().grayscale_cuda(inp, out)
    return out
scrolls · 83 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Changes from previous submission

Against this author's previous submission submission 510192.

⋯ 21 unchanged lines
#include <torch/extension.h>
#include <cuda_runtime.h>
- #define THREADS 512
+ #define THREADS 128
#define CR 0.2989f
#define CG 0.5870f
#define CB 0.1140f

Best evidence level for this revision: reported

JSON