Skip to content
KernelIndex
Search⌘K

submission 66466

shellsmile15795 · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 24 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-grayscale-v2-66466?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
RGB to grayscalesuite of 6 cases
NVIDIA B200
2.52ms
#69 of 84
2025-11-04

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:86705bf5ff096e203170f3038850e9e011a22feff442fdae3776be08d15b6a6c
license declaredunknown
license concludedunknown
authorsshellsmile15795
imported2026-08-15

Kernel source

submission.py24 lines
import torch
from task import input_t, output_t

# Standard luminance coefficients expressed as Python floats.
_WEIGHT_R = 0.2989
_WEIGHT_G = 0.5870
_WEIGHT_B = 0.1140


def custom_kernel(data: input_t) -> output_t:
    data, output = data

    # Avoid allocating temporary tensors by writing directly into the provided output buffer.
    # Using in-place arithmetic keeps the computation bandwidth-bound and reuses the output memory.
    r = data[..., 0]
    g = data[..., 1]
    b = data[..., 2]

    torch.mul(r, _WEIGHT_R, out=output)
    output.add_(g, alpha=_WEIGHT_G)
    output.add_(b, alpha=_WEIGHT_B)

    return output

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Changes from previous submission

Against this author's previous submission submission 66460.

- from task import input_t, output_t
import torch
+ from task import input_t, output_t
+ # Standard luminance coefficients expressed as Python floats.
+ _WEIGHT_R = 0.2989
+ _WEIGHT_G = 0.5870
+ _WEIGHT_B = 0.1140
+
+
def custom_kernel(data: input_t) -> output_t:
data, output = data
- weights = torch.tensor([0.2989, 0.5870, 0.1140],
- device=data.device,
- dtype=data.dtype)
- output[...] = torch.sum(data * weights, dim=-1)
+
+ # Avoid allocating temporary tensors by writing directly into the provided output buffer.
+ # Using in-place arithmetic keeps the computation bandwidth-bound and reuses the output memory.
+ r = data[..., 0]
+ g = data[..., 1]
+ b = data[..., 2]
+
+ torch.mul(r, _WEIGHT_R, out=output)
+ output.add_(g, alpha=_WEIGHT_G)
+ output.add_(b, alpha=_WEIGHT_B)
+
return output
scrolls · 28 diff lines total

Best evidence level for this revision: reported

JSON