Skip to content
KernelIndex
Search⌘K

submission 170656

Nader · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 46 lines, June 9 Researcher Reciprocity License v1.0.

histogram_v2.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-histogram-v2-170656?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesuint8

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
Histogramsuite of 6 cases
NVIDIA B200
15.1µs
#19 of 54
2025-12-17

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:2d9b7eda7bb5bb82e0efa8c054bbb8568355b36d0bab083a6c609956673850b6
license declaredunknown
license concludedunknown
authorsNader
imported2026-08-15

Kernel source

histogram_v2.py46 lines
#!POPCORN leaderboard histogram_v2

# This is a submission template for popcorn leaderboard 'histogram_v2'.
# Your task is as follows:
# > Implement a histogram kernel that counts the number of elements falling into each bin across the specified range.
# > The minimum and maximum values of the range are fixed to 0 and 100 respectively.
# > All sizes are multiples of 16 and the number of bins is set to the size of the input tensor divided by 16.
# > 
# > Input:
# >   - data: a tensor of shape (size,)
# The deadline for this leaderboard is 2025-12-30 00:00:00+00:00

# You can automatically route this file to specific GPUs by adding a line
# `#!POPCORN gpus <GPUs>` to the header of this file.
# Happy hacking!

import subprocess

subprocess.run(["pip", "install", "cuda-cccl[cu12]==0.4.3"])

from task import input_t, output_t

import cuda.compute
import numpy as np
import torch


input_size = 10485760
num_output_levels = np.array([257], dtype=np.int32)
lower_level = np.array([0], dtype=np.int32)
upper_level = np.array([256], dtype=np.int32)
build_data = torch.empty((input_size,), dtype=torch.uint8, device="cuda")
build_histogram = torch.empty((num_output_levels[0] - 1,), dtype=torch.int32, device="cuda")

histogrammer = cuda.compute.make_histogram_even(build_data, build_histogram, num_output_levels, lower_level, upper_level, input_size)

temp_storage_size = histogrammer(None, build_data, build_histogram, num_output_levels, lower_level, upper_level, input_size)
d_temp_storage = torch.empty(temp_storage_size, dtype=torch.uint8, device="cuda")

def custom_kernel(data: input_t) -> output_t:
    d_in, _ = data
    histogrammer(d_temp_storage, d_in, build_histogram, num_output_levels, lower_level, upper_level, len(d_in))

    return build_histogram

scrolls · 46 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON