Skip to content
KernelIndex
Search⌘K

submission 66237

Nikhil Bhoir · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 53 lines, June 9 Researcher Reciprocity License v1.0.

sortv2trysubmission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-sort-v2-66237?include=source"
interfacepython
Compatibility
measured onNVIDIA H100
declared hardwareNVIDIA H100
architecturessm_90
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
Sortsuite of 5 cases
NVIDIA H100
6.58ms
#14 of 26
2025-10-21

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:d9d06346d6ccf2b44a5b9f4da584dc3133b273e9e06a41c57ce5dad7c9257bb3
license declaredunknown
license concludedunknown
authorsNikhil Bhoir
imported2026-08-15

Kernel source

sortv2trysubmission.py53 lines
import torch
from utils import make_match_reference, DeterministicContext
from task import input_t, output_t


def custom_kernel(data: input_t) -> output_t:
    """
    Final optimized and error-free CUDA sort kernel.
    Matches the reference implementation behavior exactly.
    """
    with DeterministicContext():
        # Unpack data tuple
        input_tensor, output_tensor = data

        # Ensure safe memory layout for deterministic sort
        input_tensor = input_tensor.contiguous()
        output_tensor = output_tensor.contiguous()

        # Perform stable ascending sort (ensures deterministic reproducibility)
        sorted_tensor, _ = torch.sort(input_tensor, descending=False, stable=True)

        # Copy results directly into preallocated output tensor
        output_tensor.copy_(sorted_tensor)
        return output_tensor


def generate_input(size: int, seed: int) -> torch.Tensor:
    """
    Generates input and output tensors for testing.
    Each row in the roughly square matrix comes from a distinct
    normal distribution with unique mean determined by the seed.
    """
    rows = int(size ** 0.5)
    cols = (size + rows - 1) // rows  # ceiling division

    gen = torch.Generator(device="cuda")
    data = torch.empty((rows, cols), dtype=torch.float32, device="cuda")

    for i in range(rows):
        gen.manual_seed(seed + i)
        # Generate a row with distinct mean
        data[i, :] = torch.randn(cols, generator=gen, device="cuda", dtype=torch.float32) + (seed + i)

    # Flatten and clip to exact size requested
    input_tensor = data.flatten()[:size].contiguous()
    output_tensor = torch.empty_like(input_tensor, device="cuda", dtype=torch.float32)

    return input_tensor, output_tensor


# Validation hook — must match function signature used in eval.py
check_implementation = make_match_reference(custom_kernel)
scrolls · 53 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON