submission 545195
rajesh0042 · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 31 lines, June 9 Researcher Reciprocity License v1.0.
histogram_v2.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-histogram-v2-545195?include=source"interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesuint8
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:146e271fb71323dadf04b8a8119a1ac733bc69aca534de437b97e26305f8edb5
license declaredunknown
license concludedunknown
authorsrajesh0042
imported2026-08-15
Kernel source
histogram_v2.py31 lines
import os
os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"
import torch
import triton
import triton.language as tl
from task import input_t, output_t
@triton.jit
def histogram_kernel(
data_ptr, output_ptr, n_elements,
BLOCK_SIZE: tl.constexpr,
):
pid = tl.program_id(0)
offsets = pid * BLOCK_SIZE + tl.arange(0, BLOCK_SIZE)
mask = offsets < n_elements
vals = tl.load(data_ptr + offsets, mask=mask, other=0)
# Convert to int32 for indexing
vals = vals.to(tl.int32)
# Atomic add to global histogram
for i in range(BLOCK_SIZE):
if pid * BLOCK_SIZE + i < n_elements:
bin_idx = tl.load(data_ptr + pid * BLOCK_SIZE + i).to(tl.int32)
tl.atomic_add(output_ptr + bin_idx, 1)
def custom_kernel(data: input_t) -> output_t:
data, output = data
# torch.bincount is already very fast, let's just use it
output[...] = torch.bincount(data, minlength=256)
return output
scrolls · 31 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Changes from previous submission
Against this author's previous submission submission 545076.
⋯ 1 unchanged linesos.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"import torch+ import triton+ import triton.language as tlfrom task import input_t, output_t+ @triton.jit+ def histogram_kernel(+ data_ptr, output_ptr, n_elements,+ BLOCK_SIZE: tl.constexpr,+ ):+ pid = tl.program_id(0)+ offsets = pid * BLOCK_SIZE + tl.arange(0, BLOCK_SIZE)+ mask = offsets < n_elements+ vals = tl.load(data_ptr + offsets, mask=mask, other=0)+ # Convert to int32 for indexing+ vals = vals.to(tl.int32)+ # Atomic add to global histogram+ for i in range(BLOCK_SIZE):+ if pid * BLOCK_SIZE + i < n_elements:+ bin_idx = tl.load(data_ptr + pid * BLOCK_SIZE + i).to(tl.int32)+ tl.atomic_add(output_ptr + bin_idx, 1)+def custom_kernel(data: input_t) -> output_t:data, output = data+ # torch.bincount is already very fast, let's just use itoutput[...] = torch.bincount(data, minlength=256)return output
scrolls · 30 diff lines total
Best evidence level for this revision: reported
JSON