Skip to content
KernelIndex
Search⌘K

submission 512758

mreso · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 41 lines, June 9 Researcher Reciprocity License v1.0.

submission_histogram_v2.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-histogram-v2-512758?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesuint8

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
Histogramsuite of 6 cases
NVIDIA B200
1.64ms
#53 of 54
2026-03-04

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:405fef5eb0aa56fcdfdc41b38693db08ca60f5c24be1d6c36b277e24d8157fd8
license declaredunknown
license concludedunknown
authorsmreso
imported2026-08-15

Kernel source

submission_histogram_v2.py41 lines
# submission_histogram_v2.py
# 256-bin histogram of a uint8 tensor.
# Interface: custom_kernel((input_tensor, output_tensor)) -> output
#   input_tensor:  (size,) uint8  on CUDA
#   output_tensor: (256,)  int64  pre-allocated (zero-initialised here)
#
# Strategy: each thread reads one element and atomically increments its bin.
# BLOCK is fixed (not autotuned) because @triton.autotune runs the kernel
# multiple times per call; since we use global atomic_add, each trial run
# would accumulate on top of the previous one, corrupting the result.

import torch
import triton
import triton.language as tl
from task import input_t, output_t

_BLOCK = 1024
_WARPS = 4


@triton.jit
def _histogram_kernel(x_ptr, out_ptr, N: int, BLOCK: tl.constexpr):
    pid  = tl.program_id(0)
    offs = pid * BLOCK + tl.arange(0, BLOCK)
    mask = offs < N
    # Load uint8 values; cast to int64 to use as bin indices and add-type
    vals = tl.load(x_ptr + offs, mask=mask, other=0).to(tl.int64)
    # Scatter-add 1 to the corresponding bin in the global int64 histogram
    tl.atomic_add(out_ptr + vals, tl.full([BLOCK], 1, dtype=tl.int64), mask=mask)


def custom_kernel(data: input_t) -> output_t:
    x, out = data
    N = x.numel()
    out.zero_()
    if N == 0:
        return out
    P = triton.cdiv(N, _BLOCK)
    _histogram_kernel[(P,)](x, out, N, BLOCK=_BLOCK, num_warps=_WARPS)
    return out
scrolls · 41 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON