submission 545306
rajesh0042 · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 29 lines, June 9 Researcher Reciprocity License v1.0.
sort_v5.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-sort-v2-545306?include=source"interfacepython
Compatibility
measured onNVIDIA H100
declared hardwareNVIDIA H100
architecturessm_90
dtypesfp32
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:3ed966ff447fa6b97f64266e3dda7757f733ccb834a1f4163b4ad0dc22902ad8
license declaredunknown
license concludedunknown
authorsrajesh0042
imported2026-08-15
Kernel source
sort_v5.py29 lines
import os
os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"
import torch
from task import input_t, output_t
# Sort with stable=False for potential speedup
# The reference uses torch.sort(data)[0] which defaults to stable=False
# Pre-allocate both values and indices buffers
_vals_cache = {}
_idx_cache = {}
def _warmup():
for size in [1024, 4096, 16384, 65536, 262144]:
x = torch.randn(size, device='cuda', dtype=torch.float32)
torch.sort(x, stable=False)
torch.cuda.synchronize()
_warmup()
def custom_kernel(data: input_t) -> output_t:
data, output = data
n = data.numel()
if n not in _idx_cache or _idx_cache[n].device != data.device:
_idx_cache[n] = torch.empty(n, device=data.device, dtype=torch.int64)
torch.sort(data, stable=False, out=(output, _idx_cache[n]))
return output
scrolls · 29 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Changes from previous submission
Against this author's previous submission submission 545264.
⋯ 3 unchanged linesimport torchfrom task import input_t, output_t- # Pre-allocate index buffer for sort+ # Sort with stable=False for potential speedup+ # The reference uses torch.sort(data)[0] which defaults to stable=False+ # Pre-allocate both values and indices buffers++ _vals_cache = {}_idx_cache = {}+ def _warmup():+ for size in [1024, 4096, 16384, 65536, 262144]:+ x = torch.randn(size, device='cuda', dtype=torch.float32)+ torch.sort(x, stable=False)+ torch.cuda.synchronize()++ _warmup()+def custom_kernel(data: input_t) -> output_t:data, output = datan = data.numel()if n not in _idx_cache or _idx_cache[n].device != data.device:_idx_cache[n] = torch.empty(n, device=data.device, dtype=torch.int64)- torch.sort(data, out=(output, _idx_cache[n]))+ torch.sort(data, stable=False, out=(output, _idx_cache[n]))return output
scrolls · 28 diff lines total
Best evidence level for this revision: reported
JSON