submission 545313
rajesh0042 · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 29 lines, June 9 Researcher Reciprocity License v1.0.
sort_v5.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-sort-v2-545313?include=source"interfacepython
Compatibility
measured onNVIDIA L4
declared hardwareNVIDIA L4
architecturessm_89
dtypesfp32
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:f369c7725895de28323eb24e74e73ee4a597a4e514e01e38631dde73176381cb
license declaredunknown
license concludedunknown
authorsrajesh0042
imported2026-08-15
Kernel source
sort_v5.py29 lines
import os
os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"
import torch
from task import input_t, output_t
# Sort with stable=False for potential speedup
# The reference uses torch.sort(data)[0] which defaults to stable=False
# Pre-allocate both values and indices buffers
_vals_cache = {}
_idx_cache = {}
def _warmup():
for size in [1024, 4096, 16384, 65536, 262144]:
x = torch.randn(size, device='cuda', dtype=torch.float32)
torch.sort(x, stable=False)
torch.cuda.synchronize()
_warmup()
def custom_kernel(data: input_t) -> output_t:
data, output = data
n = data.numel()
if n not in _idx_cache or _idx_cache[n].device != data.device:
_idx_cache[n] = torch.empty(n, device=data.device, dtype=torch.int64)
torch.sort(data, stable=False, out=(output, _idx_cache[n]))
return output
scrolls · 29 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Changes from previous submission
Against this author's previous submission submission 545306.
Best evidence level for this revision: reported
JSON