Skip to content
KernelIndex
Search⌘K

submission 780284

Kernel-Zhang · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 63 lines, June 9 Researcher Reciprocity License v1.0.

ref.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-histogram-v2-780284?include=source"
interfacepython
Compatibility
measured onNVIDIA A100
declared hardwareNVIDIA A100
architecturessm_80
dtypesuint8

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
Histogramsuite of 6 cases
NVIDIA A100
319.6µs
#19 of 24
2026-04-27

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:5e34093e75776ed901a5bc4b68e89ca11b88b7ddda8c6e84baede1adf316e30a
license declaredunknown
license concludedunknown
authorsKernel-Zhang
imported2026-08-15

Kernel source

ref.py63 lines
from utils import verbose_allequal, DeterministicContext
import torch
from task import input_t, output_t

def custom_kernel(data: input_t) -> output_t:
    data, output = data
    # Count values in each bin
    output[...] = torch.bincount(data, minlength=256)
    return output

def ref_kernel(data: input_t) -> output_t:
    """
    Reference implementation of histogram using PyTorch.
    Args:
        data: tensor of shape (size,)
    Returns:
        Tensor containing bin counts
    """
    with DeterministicContext():
        data, output = data
        # Count values in each bin
        output[...] = torch.bincount(data, minlength=256)
        return output


def generate_input(size: int, contention: float, seed: int) -> input_t:
    """
    Generates random input tensor for histogram.

    Args:
        size: Size of the input tensor (must be multiple of 16)
        contention: float in [0, 100], specifying the percentage of identical values
        seed: Random seed
    Returns:
        The input tensor with values in [0, 255]
    """
    gen = torch.Generator(device='cuda')
    gen.manual_seed(seed)
    
    # Generate integer values between 0 and 256
    data = torch.randint(0, 256, (size,), device='cuda', dtype=torch.uint8, generator=gen)

    # make one value appear quite often, increasing the chance for atomic contention
    evil_value = torch.randint(0, 256, (), device='cuda', dtype=torch.uint8, generator=gen)
    evil_loc = torch.rand((size,), device='cuda', dtype=torch.float32, generator=gen) < (contention / 100.0)
    data[evil_loc] = evil_value

    output = torch.empty(256, device='cuda', dtype=torch.int64).contiguous()

    return data.contiguous(), output


def check_implementation(data, output):
    expected = ref_kernel(data)
    reasons = verbose_allequal(output, expected)

    if len(reasons) > 0:
        return False, "mismatch found! custom implementation doesn't match reference: " + " ".join(reasons)

    return True, ''


scrolls · 63 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON