Skip to content
KernelIndex
Search⌘K

submission 66580

albert9823 · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 18 lines, June 9 Researcher Reciprocity License v1.0.

submission_v2.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-grayscale-v2-66580?include=source"
interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
RGB to grayscalesuite of 6 cases
NVIDIA A100
8.81ms
#53 of 137
2025-11-04

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:1a834f6047ee6c24c05c1be8885b161d41c4e18e4c8da8d93e344220a818000a
license declaredunknown
license concludedunknown
authorsalbert9823
imported2026-08-15

Kernel source

submission_v2.py18 lines
from task import input_t, output_t
import torch

def custom_kernel(data: input_t) -> output_t:
    x, out = data  # x: (H, W, 3), out: (H, W)

    # out = 0.2989 * R
    torch.mul(x[..., 0], 0.2989, out=out)

    # out += 0.5870 * G
    torch.add(out, x[..., 1], alpha=0.5870, out=out)

    # out += 0.1140 * B
    torch.add(out, x[..., 2], alpha=0.1140, out=out)

    return out

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Changes from previous submission

Against this author's previous submission submission 66578.

⋯ 1 unchanged lines
import torch
def custom_kernel(data: input_t) -> output_t:
- # Unpack (input_rgb, output_gray)
x, out = data # x: (H, W, 3), out: (H, W)
- # Match device & dtype with input
- weights = torch.tensor(
- [0.2989, 0.5870, 0.1140],
- device=x.device,
- dtype=x.dtype
- )
+ # out = 0.2989 * R
+ torch.mul(x[..., 0], 0.2989, out=out)
- # Compute grayscale along the channel dim
- out[...] = torch.sum(x * weights, dim=-1)
+ # out += 0.5870 * G
+ torch.add(out, x[..., 1], alpha=0.5870, out=out)
+ # out += 0.1140 * B
+ torch.add(out, x[..., 2], alpha=0.1140, out=out)
+
return out
+
scrolls · 26 diff lines total

Best evidence level for this revision: reported

JSON