submission 66237
Nikhil Bhoir · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 53 lines, June 9 Researcher Reciprocity License v1.0.
sortv2trysubmission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-sort-v2-66237?include=source"interfacepython
Compatibility
measured onNVIDIA H100
declared hardwareNVIDIA H100
architecturessm_90
dtypesfp32
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:d9d06346d6ccf2b44a5b9f4da584dc3133b273e9e06a41c57ce5dad7c9257bb3
license declaredunknown
license concludedunknown
authorsNikhil Bhoir
imported2026-08-15
Kernel source
sortv2trysubmission.py53 lines
import torch
from utils import make_match_reference, DeterministicContext
from task import input_t, output_t
def custom_kernel(data: input_t) -> output_t:
"""
Final optimized and error-free CUDA sort kernel.
Matches the reference implementation behavior exactly.
"""
with DeterministicContext():
# Unpack data tuple
input_tensor, output_tensor = data
# Ensure safe memory layout for deterministic sort
input_tensor = input_tensor.contiguous()
output_tensor = output_tensor.contiguous()
# Perform stable ascending sort (ensures deterministic reproducibility)
sorted_tensor, _ = torch.sort(input_tensor, descending=False, stable=True)
# Copy results directly into preallocated output tensor
output_tensor.copy_(sorted_tensor)
return output_tensor
def generate_input(size: int, seed: int) -> torch.Tensor:
"""
Generates input and output tensors for testing.
Each row in the roughly square matrix comes from a distinct
normal distribution with unique mean determined by the seed.
"""
rows = int(size ** 0.5)
cols = (size + rows - 1) // rows # ceiling division
gen = torch.Generator(device="cuda")
data = torch.empty((rows, cols), dtype=torch.float32, device="cuda")
for i in range(rows):
gen.manual_seed(seed + i)
# Generate a row with distinct mean
data[i, :] = torch.randn(cols, generator=gen, device="cuda", dtype=torch.float32) + (seed + i)
# Flatten and clip to exact size requested
input_tensor = data.flatten()[:size].contiguous()
output_tensor = torch.empty_like(input_tensor, device="cuda", dtype=torch.float32)
return input_tensor, output_tensor
# Validation hook — must match function signature used in eval.py
check_implementation = make_match_reference(custom_kernel)
scrolls · 53 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON