submission 510189
ağaç.mp4 · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 83 lines, June 9 Researcher Reciprocity License v1.0.
submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-grayscale-v2-510189?include=source"interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp32
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Operation / workload
Hardware
Latency
Rank
Observed
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:5bb0aa1344f21543fe833598e4f224f94df8b5d00561f9199d571fa260e39441
license declaredunknown
license concludedunknown
authorsağaç.mp4
imported2026-08-15
Techniques
Extracted from the mirrored source by pattern, never inferred. Each row cites its line.
vector-width = float4
v6: 4px/thread, 256 threads, float4 loads+stores, explicit FMAKernel source
submission.py83 lines
"""
Grayscale — RGB to Grayscale conversion
@MemoryCoalesced
Y = 0.2989 R + 0.5870 G + 0.1140 B
Input: (H, W, 3) float32
Output: (H, W) float32
v6: 4px/thread, 256 threads, float4 loads+stores, explicit FMA
22 registers, 0 spills — 89.8% bandwidth efficiency on B200
"""
import torch
from torch.utils.cpp_extension import load_inline
cpp_source = """
void grayscale_cuda(torch::Tensor input, torch::Tensor output);
"""
cuda_source = r"""
#include <torch/extension.h>
#include <cuda_runtime.h>
#define THREADS 512
#define CR 0.2989f
#define CG 0.5870f
#define CB 0.1140f
__global__ void __launch_bounds__(THREADS)
grayscale_kernel(
const float4* __restrict__ input,
float4* __restrict__ output,
int num_groups)
{
const int idx = blockIdx.x * THREADS + threadIdx.x;
if (idx >= num_groups) return;
const float4 a = __ldg(&input[idx * 3 + 0]); // R0 G0 B0 R1
const float4 b = __ldg(&input[idx * 3 + 1]); // G1 B1 R2 G2
const float4 c = __ldg(&input[idx * 3 + 2]); // B2 R3 G3 B3
float4 result;
result.x = __fmaf_rn(CR, a.x, __fmaf_rn(CG, a.y, CB * a.z));
result.y = __fmaf_rn(CR, a.w, __fmaf_rn(CG, b.x, CB * b.y));
result.z = __fmaf_rn(CR, b.z, __fmaf_rn(CG, b.w, CB * c.x));
result.w = __fmaf_rn(CR, c.y, __fmaf_rn(CG, c.z, CB * c.w));
output[idx] = result;
}
void grayscale_cuda(torch::Tensor input, torch::Tensor output) {
const int num_pixels = input.size(0) * input.size(1);
const int num_groups = num_pixels / 4;
const int blocks = (num_groups + THREADS - 1) / THREADS;
grayscale_kernel<<<blocks, THREADS>>>(
reinterpret_cast<const float4*>(input.data_ptr<float>()),
reinterpret_cast<float4*>(output.data_ptr<float>()),
num_groups);
}
"""
module = None
def get_module():
global module
if module is None:
module = load_inline(
name='grayscale_v9',
cpp_sources=cpp_source,
cuda_sources=cuda_source,
functions=['grayscale_cuda'],
verbose=True,
extra_cuda_cflags=['-O3', '--use_fast_math', '-std=c++17'],
)
return module
def custom_kernel(data: tuple) -> torch.Tensor:
inp, out = data
get_module().grayscale_cuda(inp, out)
return out
scrolls · 83 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON