Skip to content
KernelIndex
Search⌘K

submission 68310

wecu · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 59 lines, June 9 Researcher Reciprocity License v1.0.

submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-histogram-v2-68310?include=source"
interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesuint8

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
Histogramsuite of 6 cases
NVIDIA A100
3.53ms
#24 of 24
2025-11-08

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:0aaf005d87cddd3a0f4981966427e5b073f9767eb1d9956f8de9a027cb4bec31
license declaredunknown
license concludedunknown
authorswecu
imported2026-08-15

Kernel source

submission.py59 lines
import torch
from torch.utils.cpp_extension import load_inline
from task import input_t, output_t

cpp_src = """
#include <torch/extension.h>

void histogram_kernel(torch::Tensor input, torch::Tensor output);
"""

cuda_src = """
#include <torch/extension.h>

constexpr int threadSize = 1;
constexpr int threadsPerBlock = 128;

constexpr int blockSize = threadsPerBlock * threadSize;

__global__ void histogram_block(uint8_t* input, int64_t* output, int n) {
    // Kernel implementation here

    int idx = blockIdx.x * blockDim.x + threadIdx.x;

    if (idx < n) {
        uint8_t item = input[idx];

        atomicAdd(reinterpret_cast<unsigned long long int*>(&output[item]), 1);
    } 
}

void histogram_kernel(torch::Tensor input, torch::Tensor output) {
    int n = input.numel();

    cudaMemset(output.data_ptr<int64_t>(), 0, sizeof(int64_t) * 256);
    
    histogram_block<<<(n + blockSize - 1) / blockSize, threadsPerBlock>>>(
        input.data_ptr<uint8_t>(), 
        output.data_ptr<int64_t>(), 
        n
    );
}
"""

module = load_inline(
    name="histogram_kernel_module",
    cpp_sources=cpp_src,
    cuda_sources=cuda_src,
    functions=["histogram_kernel"],
    verbose=False
)

# Note: input/output are GPU tensors!
def custom_kernel(input: input_t) -> output_t:
    inp_t, out_t = input  # Unpack the input tuple
    
    # Call the kernel directly
    module.histogram_kernel(inp_t, out_t)
    
    return out_t
scrolls · 59 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON