Skip to content
KernelIndex
Search⌘K

submission 512858

JordanNanos · python · License unknown

Use it

Vendorable · source mirrored · license unknownView source →

No package. Vendor the mirrored source: 51 lines, June 9 Researcher Reciprocity License v1.0.

sort_py_submission.py
curl "https://kernelindex.com/api/v1/implementations/kernelbot-sort-v2-512858?include=source"
interfacepython
Compatibility
measured onNVIDIA B200
declared hardwareNVIDIA B200
architecturessm_100
dtypesfp32

Benchmark evidence

1 measurement across 1 GPU, fastest first.

Operation / workload
Hardware
Latency
Rank
Observed
Sortsuite of 5 cases
NVIDIA B200
5.64ms
#19 of 23
2026-03-05

Reported · How evidence levels are derived →

Source and license

sourceavailable
revision digestsha256:e322741cfcc3a7f44eeb6238a7212b5a4cc8fea33199f974560322c6523d2142
license declaredunknown
license concludedunknown
authorsJordanNanos
imported2026-08-15

Kernel source

sort_py_submission.py51 lines
import torch
import triton
import triton.language as tl
from task import input_t, output_t


@triton.jit
def bitonic_sort_kernel(input_ptr, output_ptr, n: tl.constexpr, BLOCK_SIZE: tl.constexpr):
    offsets = tl.arange(0, BLOCK_SIZE)
    mask = offsets < n
    x = tl.load(input_ptr + offsets, mask=mask, other=float('inf'))
    x = tl.sort(x)
    tl.store(output_ptr + offsets, x, mask=mask)


def _next_power_of_two(n):
    p = 1
    while p < n:
        p <<= 1
    return p


TRITON_THRESHOLD = 65536
_sort_compiled = torch.compile(lambda x: torch.sort(x, stable=False)[0], mode="reduce-overhead")


def custom_kernel(data: input_t) -> output_t:
    inp, output = data
    n = inp.numel()

    if n <= TRITON_THRESHOLD:
        block_size = _next_power_of_two(n)
        if block_size <= 1024:
            bitonic_sort_kernel[(1,)](inp, output, n, BLOCK_SIZE=1024)
        elif block_size <= 2048:
            bitonic_sort_kernel[(1,)](inp, output, n, BLOCK_SIZE=2048)
        elif block_size <= 4096:
            bitonic_sort_kernel[(1,)](inp, output, n, BLOCK_SIZE=4096)
        elif block_size <= 8192:
            bitonic_sort_kernel[(1,)](inp, output, n, BLOCK_SIZE=8192)
        elif block_size <= 16384:
            bitonic_sort_kernel[(1,)](inp, output, n, BLOCK_SIZE=16384)
        elif block_size <= 32768:
            bitonic_sort_kernel[(1,)](inp, output, n, BLOCK_SIZE=32768)
        else:
            bitonic_sort_kernel[(1,)](inp, output, n, BLOCK_SIZE=65536)
    else:
        output.copy_(torch.sort(inp, stable=False)[0])

    return output
scrolls · 51 lines total

Source code from GPU Mode and the KernelBot dataset · June 9 Researcher Reciprocity License v1.0

Best evidence level for this revision: reported

JSON