submission 545265
rajesh0042 · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 31 lines, June 9 Researcher Reciprocity License v1.0.
histogram_v2.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-histogram-v2-545265?include=source"interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesuint8
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:6d34c7037c70f5ff87cffa0c68bea5efcef14942679ba3a495f84f557fed8b9e
license declaredunknown
license concludedunknown
authorsrajesh0042
imported2026-08-15
Kernel source
histogram_v2.py31 lines
import os
os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"
import torch
import triton
import triton.language as tl
from task import input_t, output_t
@triton.jit
def histogram_kernel(
data_ptr, output_ptr, n_elements,
BLOCK_SIZE: tl.constexpr,
):
pid = tl.program_id(0)
offsets = pid * BLOCK_SIZE + tl.arange(0, BLOCK_SIZE)
mask = offsets < n_elements
vals = tl.load(data_ptr + offsets, mask=mask, other=0)
# Convert to int32 for indexing
vals = vals.to(tl.int32)
# Atomic add to global histogram
for i in range(BLOCK_SIZE):
if pid * BLOCK_SIZE + i < n_elements:
bin_idx = tl.load(data_ptr + pid * BLOCK_SIZE + i).to(tl.int32)
tl.atomic_add(output_ptr + bin_idx, 1)
def custom_kernel(data: input_t) -> output_t:
data, output = data
# torch.bincount is already very fast, let's just use it
output[...] = torch.bincount(data, minlength=256)
return output
scrolls · 31 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Changes from previous submission
Against this author's previous submission submission 545195.
Best evidence level for this revision: reported
JSON