submission 512771
mreso · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 41 lines, June 9 Researcher Reciprocity License v1.0.
submission_histogram_v2.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-histogram-v2-512771?include=source"interfacepython
Compatibility
measured onNVIDIA L4
declared hardwareNVIDIA L4
architecturessm_89
dtypesuint8
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:08f258cb3d8ba732c3df4c456484e90342fef3ab731dc2972b1613d2b0c52fd3
license declaredunknown
license concludedunknown
authorsmreso
imported2026-08-15
Kernel source
submission_histogram_v2.py41 lines
# submission_histogram_v2.py
# 256-bin histogram of a uint8 tensor.
# Interface: custom_kernel((input_tensor, output_tensor)) -> output
# input_tensor: (size,) uint8 on CUDA
# output_tensor: (256,) int64 pre-allocated (zero-initialised here)
#
# Strategy: each thread reads one element and atomically increments its bin.
# BLOCK is fixed (not autotuned) because @triton.autotune runs the kernel
# multiple times per call; since we use global atomic_add, each trial run
# would accumulate on top of the previous one, corrupting the result.
import torch
import triton
import triton.language as tl
from task import input_t, output_t
_BLOCK = 1024
_WARPS = 4
@triton.jit
def _histogram_kernel(x_ptr, out_ptr, N: int, BLOCK: tl.constexpr):
pid = tl.program_id(0)
offs = pid * BLOCK + tl.arange(0, BLOCK)
mask = offs < N
# Load uint8 values; cast to int64 to use as bin indices and add-type
vals = tl.load(x_ptr + offs, mask=mask, other=0).to(tl.int64)
# Scatter-add 1 to the corresponding bin in the global int64 histogram
tl.atomic_add(out_ptr + vals, tl.full([BLOCK], 1, dtype=tl.int64), mask=mask)
def custom_kernel(data: input_t) -> output_t:
x, out = data
N = x.numel()
out.zero_()
if N == 0:
return out
P = triton.cdiv(N, _BLOCK)
_histogram_kernel[(P,)](x, out, N, BLOCK=_BLOCK, num_warps=_WARPS)
return out
scrolls · 41 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Changes from previous submission
Against this author's previous submission submission 512758.
Best evidence level for this revision: reported
JSON