submission 512858
JordanNanos · python · License unknown
Use it
Vendorable · source mirrored · license unknownView source →
No package. Vendor the mirrored source: 51 lines, June 9 Researcher Reciprocity License v1.0.
sort_py_submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-sort-v2-512858?include=source"interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp32
Benchmark evidence
1 measurement across 1 GPU, fastest first.
Reported · How evidence levels are derived →
Source and license
sourceavailable
revision digestsha256:e322741cfcc3a7f44eeb6238a7212b5a4cc8fea33199f974560322c6523d2142
license declaredunknown
license concludedunknown
authorsJordanNanos
imported2026-08-15
Kernel source
sort_py_submission.py51 lines
import torch
import triton
import triton.language as tl
from task import input_t, output_t
@triton.jit
def bitonic_sort_kernel(input_ptr, output_ptr, n: tl.constexpr, BLOCK_SIZE: tl.constexpr):
offsets = tl.arange(0, BLOCK_SIZE)
mask = offsets < n
x = tl.load(input_ptr + offsets, mask=mask, other=float('inf'))
x = tl.sort(x)
tl.store(output_ptr + offsets, x, mask=mask)
def _next_power_of_two(n):
p = 1
while p < n:
p <<= 1
return p
TRITON_THRESHOLD = 65536
_sort_compiled = torch.compile(lambda x: torch.sort(x, stable=False)[0], mode="reduce-overhead")
def custom_kernel(data: input_t) -> output_t:
inp, output = data
n = inp.numel()
if n <= TRITON_THRESHOLD:
block_size = _next_power_of_two(n)
if block_size <= 1024:
bitonic_sort_kernel[(1,)](inp, output, n, BLOCK_SIZE=1024)
elif block_size <= 2048:
bitonic_sort_kernel[(1,)](inp, output, n, BLOCK_SIZE=2048)
elif block_size <= 4096:
bitonic_sort_kernel[(1,)](inp, output, n, BLOCK_SIZE=4096)
elif block_size <= 8192:
bitonic_sort_kernel[(1,)](inp, output, n, BLOCK_SIZE=8192)
elif block_size <= 16384:
bitonic_sort_kernel[(1,)](inp, output, n, BLOCK_SIZE=16384)
elif block_size <= 32768:
bitonic_sort_kernel[(1,)](inp, output, n, BLOCK_SIZE=32768)
else:
bitonic_sort_kernel[(1,)](inp, output, n, BLOCK_SIZE=65536)
else:
output.copy_(torch.sort(inp, stable=False)[0])
return output
scrolls · 51 lines total
Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0
Best evidence level for this revision: reported
JSON